AWS Certified Machine Learning – SpecialtyModelingEasy
A data science team is developing a machine learning model to predict customer lifetime value (CLV). They have collected historical data including customer demographics, purchase history, and engagement metrics. The initial model, a linear regression, shows a high bias, consistently underestimating the CLV for high-value customers and overestimating for low-value customers. Which of the following strategies is most likely to address this high bias and improve model accuracy?
- ASwitch to a more complex model like a Gradient Boosting Machine or a Neural Network.
- BApply L1 regularization (Lasso) to the linear regression model.
- CReduce the number of features used in the linear regression model.
- DIncrease the amount of training data by collecting more historical records.
Show answer & explanationAnswer & explanation
Correct answer: A. Switch to a more complex model like a Gradient Boosting Machine or a Neural Network.
High bias indicates that the model is too simple to capture the underlying patterns in the data. Switching to a more complex model can increase model capacity and reduce bias.
Why the other options are wrong
- B. L1 regularization simplifies the model by pushing some coefficients to zero, which would likely increase bias further.
- C. Reducing features simplifies the model, which would likely increase bias, making it less capable of learning complex relationships.
- D. Increasing data alone won't fix a fundamentally underfit model; it might exacerbate the issue if the model can't learn from it.
Bias-Variance Trade-off
The dilemma in supervised learning where reducing one type of error (bias or variance) often increases the other.
- High bias (underfitting) means the model is too simple.
- High variance (overfitting) means the model is too complex.
- The goal is to find a balance for optimal generalization.
Memory trick: Simple models are biased, complex models are varied.