A financial institution is implementing a machine learning model to detect fraudulent transactions. The model is designed to flag suspicious activities, and the primary concern is to minimize the number of legitimate transactions that are incorrectly flagged as fraudulent, as this can lead to customer dissatisfaction and operational overhead. Which evaluation metric should the institution prioritize to ensure this objective is met?
- ARecall
- BPrecision
- CF1-Score
- DAccuracy
Show answer & explanationAnswer & explanation
Correct answer: B. Precision
The institution's primary concern is to minimize 'legitimate transactions that are incorrectly flagged as fraudulent'. This directly relates to 'false positives' (non-fraudulent transactions incorrectly identified as fraud). Precision measures the proportion of positive identifications (flagged as fraud) that were actually correct. High precision means fewer false positives, directly addressing the goal of reducing customer dissatisfaction and operational overhead from incorrect flags.
Why the other options are wrong
- A. Recall (Sensitivity) measures the proportion of actual positive cases that are correctly identified, focusing on minimizing false negatives (missed fraud).
- C. F1-Score is the harmonic mean of Precision and Recall, providing a balance, but does not prioritize minimizing false positives specifically.
- D. Accuracy measures the overall correctness of the model, which can be misleading in imbalanced datasets like fraud detection.
Precision
A metric that measures the proportion of positive identifications that were actually correct. It is calculated as True Positives / (True Positives + False Positives).
- Focuses on minimizing False Positives (incorrectly identified positives).
- Important when the cost of a false alarm is high (e.g., flagging legitimate transactions as fraud, spam detection).
- Answers the question: 'Of all items predicted positive, how many are actually positive?'
Memory trick: PR-F1, Accuracy's good, but Precision avoids false alarms.