AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard

A data engineering team is building a feature store for a machine learning pipeline. They have ingested a massive dataset, and during the initial Exploratory Data Analysis (EDA), they discover that a critical numerical feature, 'customer_lifetime_value', has a long tail of extremely high values (positive skew) and also contains a significant number of zero values, representing customers with no recorded value yet. A standard logarithmic transformation (log(x)) would fail for zero values. To normalize this feature for model training while preserving its relative relationships and handling zeros, which transformation strategy is most appropriate?

  1. ALogarithmic transformation with a small constant added (log(x+c)), to handle zero values and reduce skewness.
  2. BMin-Max Scaling, as it scales the data to a fixed range (e.g., 0 to 1).
  3. CBox-Cox transformation, as it can handle non-normal distributions and automatically find the best power transformation.
  4. DStandard Scaling (Z-score normalization), as it centers the data around zero and scales to unit variance.
Show answer & explanation

Correct answer: A. Logarithmic transformation with a small constant added (log(x+c)), to handle zero values and reduce skewness.

A logarithmic transformation (log(x)) is effective for reducing positive skewness. To handle zero values, a small constant (c) is added to x before taking the logarithm (log(x+c)). This ensures that zero values become log(c), which is a valid number, and it still effectively compresses the skewed distribution.

Why the other options are wrong

  • B. Incorrect. Min-Max Scaling scales the data but does not address the skewness of the distribution, nor does it prepare the data for a direct logarithmic transformation if zeros are present.
  • C. Incorrect. While Box-Cox is a powerful power transformation for non-normal data, it specifically requires strictly positive values for all data points. It cannot directly handle zero values without prior modification or imputation, making log(x+c) a more straightforward and robust choice for this specific scenario with zeros.
  • D. Incorrect. Standard Scaling normalizes without addressing skewness or the specific issue of zero values in a log transform. It would still result in a skewed distribution.

Log Transformation with Constant

A data transformation technique (log(x+c)) used to reduce positive skewness in a numerical feature, where a small positive constant (c) is added to each value to allow transformation of zero or negative values.

  • Effective for right-skewed data.
  • Handles zero values, converting them to log(c).
  • Preserves order and relative relationships.

Memory trick: Log(x+c) makes skewed data agree, especially when zero's in the tree.

More Exploratory Data Analysis questions