CFA Level II ExamQuantitative MethodsHard
A financial institution is developing a machine learning model to predict loan defaults. The dataset is highly imbalanced, with only 5% of loans resulting in default. The initial model, a logistic regression, achieves 95% accuracy. However, upon further inspection, it is found that the model simply predicts 'no default' for all loans. Which of the following metrics would be most appropriate to evaluate the model's true performance in this scenario?
- ARecall
- BAccuracy
- CPrecision
- DF1-score
Show answer & explanationAnswer & explanation
Correct answer: D. F1-score
In highly imbalanced datasets, accuracy can be misleading (as seen, 95% accuracy by predicting no defaults). The F1-score is the harmonic mean of precision and recall, providing a balanced measure that is particularly useful when dealing with imbalanced classes. It penalizes models that perform poorly on either precision or recall, making it a robust metric for this scenario.
Why the other options are wrong
- A. Recall (True Positives / (True Positives + False Negatives)) measures the proportion of actual positive cases that were correctly identified. A model predicting 'no default' for all would have a recall of 0% for the 'default' class, which is informative but doesn't consider false positives. A single metric is often preferred for overall evaluation.
- B. Accuracy is misleading here because the model can achieve high accuracy (95%) by simply predicting the majority class ('no default') and ignoring the minority class ('default').
- C. Precision (True Positives / (True Positives + False Positives)) measures the proportion of correctly identified positive cases among all cases predicted as positive. While important, a model predicting 'no default' for all would have undefined precision (0/0) or very low precision if it incorrectly predicted any default. It alone doesn't capture the full picture.
F1-score
The F1-score is the harmonic mean of precision and recall, providing a balanced measure of a model's accuracy, particularly useful for classification problems with imbalanced datasets where both false positives and false negatives are important.
- F1 = 2 * (Precision * Recall) / (Precision + Recall)
- Ranges from 0 to 1, with 1 being perfect.
- Penalizes models with poor performance on either precision or recall.
- Especially relevant for minority class evaluation.
Memory trick: Balance your evaluation with F1.