AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisEasy
A data scientist is preparing a dataset for a machine learning model. They notice that a significant portion of a numerical feature, 'customer_age', is missing. Directly removing rows with missing values would lead to substantial data loss. Which imputation strategy is most appropriate if the distribution of 'customer_age' is heavily skewed and they want to minimize the impact of outliers on the imputed values?
- AMean imputation
- BZero imputation
- CMode imputation
- DMedian imputation
Show answer & explanationAnswer & explanation
Correct answer: D. Median imputation
Median imputation is robust to outliers and skewed distributions because it uses the middle value, which is less affected by extreme values than the mean. This helps preserve the underlying distribution characteristics better than mean imputation when skewness is present.
Why the other options are wrong
- A. Mean imputation is highly sensitive to outliers and skewed distributions, which would distort the imputed values.
- B. Zero imputation assumes that missing values genuinely represent zero, which is unlikely for 'customer_age' and would severely bias the data.
- C. Mode imputation is typically used for categorical features, not numerical features like age.
Median Imputation
A data imputation technique where missing values in a feature are replaced with the median of the observed values for that feature.
- Robust to outliers and skewed distributions.
- Preserves rank order better than mean imputation.
- Can reduce variance of the imputed variable.
Memory trick: Missing data: MEDIAN is the most 'middle-ground' for skewed data.