CFA Level II ExamQuantitative MethodsMedium
A portfolio manager is evaluating a new machine learning algorithm designed to predict stock price movements. The algorithm was trained on historical data and achieved 90% accuracy on the training set. However, when tested on a separate, unseen validation set, the accuracy dropped to 65%. This scenario is most indicative of which of the following issues?
- ABias-variance trade-off
- BOverfitting
- CData leakage
- DUnderfitting
Show answer & explanationAnswer & explanation
Correct answer: B. Overfitting
Overfitting occurs when a model learns the training data too well, including its noise and random fluctuations, leading to high accuracy on the training set but poor generalization and significantly lower accuracy on unseen data (validation or test set).
Why the other options are wrong
- A. The bias-variance trade-off is a fundamental concept in machine learning, but 'overfitting' is the specific manifestation of high variance in this scenario. It's a consequence, not the issue itself.
- C. Data leakage occurs when information from the validation or test set is inadvertently used during training, leading to overly optimistic performance estimates, but it wouldn't cause a drop in accuracy on the validation set unless the validation set was then truly unseen (i.e., the leakage was only partial or from a different source). Here, the drop is the key symptom.
- D. Underfitting occurs when a model is too simple to capture the underlying patterns in the data, resulting in poor performance on both training and validation sets.
Overfitting
Overfitting in machine learning occurs when a model learns the training data too precisely, capturing noise and specific patterns that do not generalize well to new, unseen data, resulting in high training accuracy but low validation/test accuracy.
- High variance, low bias.
- Symptoms: Good performance on training data, poor performance on unseen data.
- Mitigation: Regularization, cross-validation, more data, simpler models.
Memory trick: A model that 'over-studies' performs poorly on the real test.