Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureMedium
A data scientist is training a machine learning model. During the training process, the model's performance on the training dataset continuously improves, achieving near-perfect accuracy. However, when evaluated on a separate, unseen validation dataset, the model's performance is significantly worse. This indicates a common problem in machine learning. What is this problem called?
- AOverfitting
- BVariance
- CUnderfitting
- DBias
Show answer & explanationAnswer & explanation
Correct answer: A. Overfitting
This scenario describes overfitting, where a model learns the training data too well, including its noise and specific patterns, leading to poor generalization on new, unseen data.
Why the other options are wrong
- B. Variance refers to the amount that the estimate of the target function will change if different training data was used, often associated with overfitting.
- C. Underfitting occurs when a model is too simple to capture the underlying patterns in the training data, leading to poor performance on both training and validation sets.
- D. Bias refers to the simplifying assumptions made by a model to make the target function easier to learn, often leading to underfitting.
Overfitting
A modeling error that occurs when a function is too closely or exactly fitted to a limited set of data points. It performs well on training data but poorly on new, unseen data.
- Model learns noise and specific patterns of the training data.
- High performance on training set, low performance on validation/test set.
- Can be mitigated by regularization, more data, simpler models, cross-validation.
Memory trick: Fit the Data, Not Just the Noise