AWS Certified Machine Learning – SpecialtyModelingEasy
A data science team is developing a machine learning model to predict customer churn. They observe that the dataset is highly imbalanced, with a very small percentage of customers actually churning. The business objective is to identify as many churning customers as possible to intervene proactively, even if it means some false positives. Which evaluation metric should the team prioritize to align with this objective?
- AAccuracy
- BPrecision
- CRecall
- DF1-score
Show answer & explanationAnswer & explanation
Correct answer: C. Recall
The business objective is to identify as many churning customers as possible, which means maximizing the true positive rate. Recall (also known as sensitivity) directly measures the proportion of actual positive cases that are correctly identified by the model, making it the most suitable metric for this scenario.
Why the other options are wrong
- A. Accuracy can be misleading in imbalanced datasets as a model that predicts the majority class can still have high accuracy.
- B. Precision measures the proportion of positive identifications that were actually correct, which is less relevant if false positives are acceptable.
- D. F1-score is the harmonic mean of precision and recall, providing a balance between the two, but the objective here strongly prioritizes recall.
Recall (Sensitivity)
Recall, also known as sensitivity or true positive rate, measures the proportion of actual positive cases that are correctly identified by the model.
- Calculated as True Positives / (True Positives + False Negatives).
- Prioritized when the cost of False Negatives is high (e.g., missing a disease, fraud).
- Aims to minimize false negatives.
Memory trick: Remember, 'Recall' is like a 'fishing net' catching all the 'churners', even if some 'non-churners' get caught too.