AWS Certified Machine Learning – SpecialtyModelingEasy

A data science team is developing a machine learning model to predict customer lifetime value (CLV). They have collected historical data including customer demographics, purchase history, and engagement metrics. The initial model, a linear regression, shows a high bias, consistently underestimating the CLV for high-value customers and overestimating for low-value customers. Which of the following strategies is most likely to address this high bias and improve model accuracy?

  1. ASwitch to a more complex model like a Gradient Boosting Machine or a Neural Network.
  2. BApply L1 regularization (Lasso) to the linear regression model.
  3. CReduce the number of features used in the linear regression model.
  4. DIncrease the amount of training data by collecting more historical records.
Show answer & explanation

Correct answer: A. Switch to a more complex model like a Gradient Boosting Machine or a Neural Network.

High bias indicates that the model is too simple to capture the underlying patterns in the data. Switching to a more complex model can increase model capacity and reduce bias.

Why the other options are wrong

  • B. L1 regularization simplifies the model by pushing some coefficients to zero, which would likely increase bias further.
  • C. Reducing features simplifies the model, which would likely increase bias, making it less capable of learning complex relationships.
  • D. Increasing data alone won't fix a fundamentally underfit model; it might exacerbate the issue if the model can't learn from it.

Bias-Variance Trade-off

The dilemma in supervised learning where reducing one type of error (bias or variance) often increases the other.

  • High bias (underfitting) means the model is too simple.
  • High variance (overfitting) means the model is too complex.
  • The goal is to find a balance for optimal generalization.

Memory trick: Simple models are biased, complex models are varied.

More Modeling questions