AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard
A data engineering team is building a feature store for a machine learning pipeline. They have ingested a massive dataset, and during the initial Exploratory Data Analysis (EDA), they discover that a critical numerical feature, 'customer_lifetime_value', has a long tail of extremely high values (positive skew) and also contains a significant number of zero values, representing customers with no recorded value yet. A standard logarithmic transformation (log(x)) would fail for zero values. To normalize this feature for model training while preserving its relative relationships and handling zeros, which transformation strategy is most appropriate?
- ALogarithmic transformation with a small constant added (log(x+c)), to handle zero values and reduce skewness.
- BMin-Max Scaling, as it scales the data to a fixed range (e.g., 0 to 1).
- CBox-Cox transformation, as it can handle non-normal distributions and automatically find the best power transformation.
- DStandard Scaling (Z-score normalization), as it centers the data around zero and scales to unit variance.
Show answer & explanationAnswer & explanation
Correct answer: A. Logarithmic transformation with a small constant added (log(x+c)), to handle zero values and reduce skewness.
A logarithmic transformation (log(x)) is effective for reducing positive skewness. To handle zero values, a small constant (c) is added to x before taking the logarithm (log(x+c)). This ensures that zero values become log(c), which is a valid number, and it still effectively compresses the skewed distribution.
Why the other options are wrong
- B. Incorrect. Min-Max Scaling scales the data but does not address the skewness of the distribution, nor does it prepare the data for a direct logarithmic transformation if zeros are present.
- C. Incorrect. While Box-Cox is a powerful power transformation for non-normal data, it specifically requires strictly positive values for all data points. It cannot directly handle zero values without prior modification or imputation, making log(x+c) a more straightforward and robust choice for this specific scenario with zeros.
- D. Incorrect. Standard Scaling normalizes without addressing skewness or the specific issue of zero values in a log transform. It would still result in a skewed distribution.
Log Transformation with Constant
A data transformation technique (log(x+c)) used to reduce positive skewness in a numerical feature, where a small positive constant (c) is added to each value to allow transformation of zero or negative values.
- Effective for right-skewed data.
- Handles zero values, converting them to log(c).
- Preserves order and relative relationships.
Memory trick: Log(x+c) makes skewed data agree, especially when zero's in the tree.