AWS Certified Machine Learning – SpecialtyModelingMedium
A data scientist is evaluating a binary classification model for predicting a rare disease. The model's performance is assessed using various metrics. The business requirement prioritizes minimizing false negatives, as missing a positive case (disease present) is considered far more costly than a false positive (healthy person incorrectly diagnosed with the disease). Which metric should the data scientist focus on optimizing?
- ARecall (Sensitivity)
- BPrecision
- CF1-Score
- DAccuracy
Show answer & explanationAnswer & explanation
Correct answer: A. Recall (Sensitivity)
Recall (Sensitivity) measures the proportion of actual positive cases (disease present) that were correctly identified. In this scenario, minimizing false negatives means ensuring as many actual disease cases are detected as possible, which directly corresponds to maximizing recall.
Why the other options are wrong
- B. Precision measures the proportion of positive predictions that were actually correct. While important, optimizing precision alone might lead to missing many actual positive cases (high false negatives).
- C. F1-Score is the harmonic mean of precision and recall. While useful for an overall balance, if the primary goal is to minimize false negatives, recall is a more direct metric to optimize.
- D. Accuracy can be misleading in imbalanced datasets (rare disease) and does not specifically prioritize minimizing false negatives over other error types.
Recall (Sensitivity)
The proportion of actual positive instances that were correctly identified by the model. It quantifies the model's ability to 'find all the positive samples'.
- Calculated as True Positives / (True Positives + False Negatives).
- High recall means fewer false negatives.
- Crucial when the cost of missing a positive is high (e.g., disease detection, fraud detection).
Memory trick: When missing a positive is a terrible fright, recall's your beacon, shining ever so bright.