CompTIA Data+ (DA0-002)Data AnalysisHard
A data scientist is developing a new algorithm to detect fraudulent transactions. They train the algorithm and now need to evaluate its performance. A critical requirement is to minimize the number of actual fraudulent transactions that are missed by the algorithm, even if it means flagging some legitimate transactions as fraudulent. Which evaluation metric should the data scientist prioritize to meet this requirement?
- ASpecificity
- BRecall
- CPrecision
- DAccuracy
Show answer & explanationAnswer & explanation
Correct answer: B. Recall
Recall (also known as sensitivity or true positive rate) measures the proportion of actual positive cases (actual fraudulent transactions) that were correctly identified by the model. Prioritizing recall minimizes false negatives (missed fraudulent transactions) at the potential cost of more false positives.
Why the other options are wrong
- A. Specificity measures the proportion of actual negative cases (legitimate transactions) that were correctly identified as legitimate (minimizing true negatives, not directly related to missed fraud).
- C. Precision measures the proportion of correctly identified fraudulent transactions among all transactions flagged as fraudulent (minimizing false positives).
- D. Accuracy measures overall correct predictions but can be misleading if the classes are imbalanced (e.g., very few fraudulent transactions).
Recall (Sensitivity)
The proportion of actual positive instances that were correctly identified by the model. It focuses on minimizing false negatives.
- Formula: True Positives / (True Positives + False Negatives).
- Important when the cost of a false negative is high (e.g., fraud detection, medical diagnosis).
- Also known as the True Positive Rate (TPR).
Memory trick: Recall 'R'emembers all the real ones.