AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard

A data engineering team is building a feature store for a machine learning pipeline. They have a numerical feature, 'customer_session_duration_seconds', which is heavily right-skewed with a long tail of very high values. The machine learning model (e.g., linear regression) they plan to use performs better with normally distributed features. To reduce skewness and stabilize variance, which transformation should they apply, ensuring no issues with zero or negative values?

  1. AMin-Max Scaling
  2. BLog Transformation (log(x))
  3. CLog Transformation (log(x + C)) with a small positive constant C
  4. DStandardization (Z-score normalization)
Show answer & explanation

Correct answer: C. Log Transformation (log(x + C)) with a small positive constant C

A log transformation is effective for reducing right-skewness and stabilizing variance. However, if 'customer_session_duration_seconds' can be zero, `log(x)` would result in undefined values. By adding a small positive constant `C` (e.g., `log(x+1)`), the transformation can handle zero values gracefully while still achieving the desired effect on skewness and variance.

Why the other options are wrong

  • A. Min-Max Scaling scales data to a specific range (e.g., 0-1) but does not address skewness or stabilize variance.
  • B. A simple Log Transformation (log(x)) would result in undefined values if 'customer_session_duration_seconds' can be zero, which is a common occurrence for durations.
  • D. Standardization centers the data around zero and scales by standard deviation but does not address skewness or stabilize variance effectively for highly skewed data.

Log Transformation with Constant

A data transformation (log(x+C)) used to reduce right-skewness, stabilize variance, and improve normality in numerical features, particularly when the original data may contain zero values.

  • Effective for right-skewed data.
  • The constant C ensures the logarithm is well-defined for x=0.
  • Commonly used in preparing data for linear models.

Memory trick: Log(x+C): 'L'ong 'O'utliers 'G'one, 'C'alm 'C'urves.

More Exploratory Data Analysis questions