AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A machine learning engineer is deploying a model that classifies customer support tickets by issue type. The model achieves 98% accuracy on the training data but only 65% accuracy on new, unseen customer tickets. What is the most likely issue with this model?
- AOverfitting
- BHigh Variance in Training Data
- CUnderfitting
- DBias
Show answer & explanationAnswer & explanation
Correct answer: A. Overfitting
Overfitting occurs when a model learns the training data too well, including its noise and specific patterns, leading to excellent performance on training data but poor generalization to new, unseen data. The significant drop in accuracy from 98% on training to 65% on unseen data is a classic symptom of overfitting.
Why the other options are wrong
- B. High variance in training data itself doesn't directly cause this symptom; rather, the model's response to that variance (or lack thereof) leads to overfitting.
- C. Underfitting is when a model is too simple to capture the underlying patterns, leading to poor performance on both training and test data.
- D. Bias refers to systematic errors or assumptions in the learning algorithm, often leading to underfitting.
Overfitting
A modeling error that occurs when a function is too closely aligned to a limited set of data points, performing well on training data but poorly on unseen data.
- High accuracy on training data.
- Low accuracy on validation/test data.
- Model learns noise and specific patterns from training data.
Memory trick: Models 'fit' data, but sometimes they 'over' or 'under' do it.