AWS Certified Machine Learning – SpecialtyModelingMedium

A financial institution is developing a machine learning model to predict stock market volatility. The historical volatility data is observed to be highly skewed, with a long tail towards higher volatility values. The data scientist understands that many classic machine learning algorithms perform best when input features are normally distributed. Which data transformation technique should be applied to the volatility feature to make it more amenable to these algorithms?

  1. AOne-Hot Encoding
  2. BLog Transformation
  3. CStandardization (Z-score normalization)
  4. DMin-Max Scaling
Show answer & explanation

Correct answer: B. Log Transformation

Log transformation is particularly effective for highly skewed positive data, like volatility or financial data, that has a long tail. It compresses the range of values, especially large ones, making the distribution more symmetrical and closer to a normal distribution. This helps algorithms that assume normality or are sensitive to feature scales to perform better.

Why the other options are wrong

  • A. One-Hot Encoding is used for categorical features, not numerical features like volatility, and does not address skewness.
  • C. Standardization scales data to have a mean of 0 and standard deviation of 1 but does not fundamentally change the shape of a skewed distribution.
  • D. Min-Max Scaling scales data to a fixed range (e.g., 0-1) but does not address skewness.

Log Transformation for Skewed Data

Log transformation is a data preprocessing technique that applies the logarithm function to a numerical feature, commonly used to reduce skewness and stabilize variance in positively skewed distributions.

  • Effective for features with a long tail towards larger values (positive skew).
  • Helps make the data distribution more symmetrical, closer to normal.
  • Beneficial for algorithms sensitive to feature distribution and scale.

Memory trick: When data is 'skewed', 'log' it to 'straighten' it out!

More Modeling questions