AWS Certified Machine Learning – SpecialtyModelingMedium

A data scientist is developing a machine learning model to predict customer churn. The dataset contains various customer attributes, including demographics, service usage, and historical interactions. During exploratory data analysis, the data scientist discovers that several numerical features have a non-linear relationship with the target variable and are heavily skewed. The chosen model, a linear regression, is underperforming. Which data transformation technique would be most beneficial for improving the linear regression model's performance in this scenario?

  1. AMin-Max Scaling.
  2. BLog Transformation.
  3. CStandardization (Z-score normalization).
  4. DOne-Hot Encoding.
Show answer & explanation

Correct answer: B. Log Transformation.

Linear regression assumes a linear relationship between features and the target, and it is sensitive to skewed data and outliers. Log transformation is highly effective at reducing skewness in positively skewed data and can help to linearize non-linear relationships, making the data more suitable for linear models.

Why the other options are wrong

  • A. Min-Max scaling scales data to a specific range but also doesn't address skewness or non-linearity.
  • C. Standardization centers and scales data but doesn't address skewness or non-linearity.
  • D. One-Hot Encoding is for categorical variables, not numerical features with skewness or non-linear relationships.

Log Transformation for Skewed Data

A data transformation technique (e.g., natural log or base-10 log) applied to numerical features to reduce skewness and potentially linearize non-linear relationships, making the data more suitable for models that assume linearity or normality.

  • Effective for positively skewed data.
  • Can stabilize variance.
  • Helps linear models capture non-linear patterns.

Memory trick: Log It Down to Straighten the Line.

More Modeling questions