AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium
A data scientist is working with a highly skewed dataset of customer income, where a few customers have extremely high incomes, significantly distorting the mean. The goal is to transform this data to achieve a more symmetrical distribution, which is a common assumption for many machine learning models. Which transformation is most appropriate in this scenario?
- AMin-Max Normalization
- BLog Transformation
- CSquare Root Transformation
- DReciprocal Transformation
Show answer & explanationAnswer & explanation
Correct answer: B. Log Transformation
Log transformation (e.g., natural log or log base 10) is highly effective for reducing positive skewness in data by compressing the larger values much more than the smaller values, making the distribution more symmetrical.
Why the other options are wrong
- A. Min-Max Normalization scales data to a specific range (e.g., 0-1) but does not address skewness or distribution shape.
- C. Square Root Transformation is less aggressive than log transformation and might not sufficiently handle extreme skewness.
- D. Reciprocal Transformation (1/x) can reverse the order of values and is typically used for specific types of relationships or when skewness is very severe and log isn't enough, but it can also introduce new issues.
Log Transformation
A data transformation technique that replaces each data point with its logarithm. It is commonly used to reduce skewness and stabilize variance in positively skewed distributions.
- Effective for highly right-skewed data (e.g., income, house prices).
- Compresses large values more than small values.
- Helps achieve a more symmetrical, normal-like distribution.
- Requires all data points to be positive; add a constant if zeros or negatives exist.
Memory trick: Skewed income? Log it to make it look normal!