AWS Certified Machine Learning – SpecialtyModelingMedium

A machine learning team is developing a model to predict house prices. They have collected a dataset with various features like square footage, number of bedrooms, location, and year built. After initial training, they find that the model consistently underpredicts high-value houses and overpredicts low-value houses. This indicates a potential issue with the model's ability to capture the non-linear relationship between features and target variable, especially at the extremes. Which model training approach is most likely to resolve this issue?

  1. AApplying a Gradient Boosting Regressor (e.g., LightGBM or XGBoost).
  2. BReducing the number of features through Principal Component Analysis (PCA).
  3. CUsing a simple Linear Regression model with L1 regularization.
  4. DIncreasing the learning rate of a Stochastic Gradient Descent (SGD) optimizer.
Show answer & explanation

Correct answer: A. Applying a Gradient Boosting Regressor (e.g., LightGBM or XGBoost).

Gradient Boosting Regressors like LightGBM or XGBoost are powerful ensemble methods that build models sequentially, where each new model corrects the errors of the previous ones. This iterative error correction, especially focusing on hard-to-predict instances, makes them highly effective at capturing complex, non-linear relationships and improving predictions across the entire range of values, including extremes. They are less prone to the bias seen in simpler models.

Why the other options are wrong

  • B. PCA reduces dimensionality but does not inherently improve the model's ability to capture non-linear relationships; it might even remove useful non-linear information.
  • C. Linear Regression is inherently linear and would struggle with non-linear relationships, and L1 regularization primarily helps with feature selection, not fixing non-linearity bias.
  • D. Increasing the learning rate of SGD might speed up convergence but could also lead to instability or divergence, and does not directly address the model's fundamental inability to learn non-linear patterns.

Gradient Boosting Regressor

Gradient Boosting is a machine learning ensemble technique that builds a strong predictive model from a sequence of weaker base models (typically decision trees). Each subsequent model attempts to correct the errors of the previous models.

  • Sequentially builds models to minimize a loss function.
  • Effective at capturing complex non-linear relationships.
  • Often achieves state-of-the-art performance in tabular data.
  • Examples include XGBoost, LightGBM, CatBoost.

Memory trick: To fix 'BIASED PREDICTIONS', you need a 'SMARTER, ADAPTIVE' model.

More Modeling questions