AWS Certified Machine Learning – SpecialtyData EngineeringMedium
A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a 'customer_age' column, which has a significant number of missing values. The data scientist observes that the distribution of existing 'customer_age' values is heavily skewed towards younger customers (ages 20-35) with a long tail extending to older ages. Which imputation strategy is most appropriate to handle the missing 'customer_age' values while preserving the original distribution characteristics?
- ARegression imputation
- BMean imputation
- CMode imputation
- DMedian imputation
Show answer & explanationAnswer & explanation
Correct answer: D. Median imputation
For skewed distributions, the median is a more robust measure of central tendency than the mean, as it is less affected by outliers. Imputing with the median will help preserve the shape of the original distribution better than the mean or mode.
Why the other options are wrong
- A. Regression imputation is more complex and effective when there are strong correlations with other features, but for simple missing value handling in a skewed distribution, median is often a good first choice without additional information.
- B. Mean imputation is sensitive to outliers and skewness, which would distort the original distribution of 'customer_age'.
- C. Mode imputation is generally used for categorical data or highly discrete numerical data, not ideal for a continuous skewed distribution.
Median Imputation for Skewed Data
Replacing missing numerical values with the median of the existing values in that column, particularly effective for skewed distributions to maintain data integrity.
- Robust to outliers, unlike the mean.
- Preserves the shape of skewed distributions better.
- Suitable for continuous numerical features with missing values.
Memory trick: Missing data? Medians for skewed, Means for normal, Modes for categories!