AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisEasy

A data scientist is preparing a dataset for a machine learning model. They notice that a significant portion of the 'income' feature (approximately 15%) is missing. The distribution of the existing 'income' data is heavily skewed to the right, with a long tail of high values. Simple mean imputation would distort the distribution. Which imputation strategy is most appropriate to maintain the integrity of the data distribution while addressing the missing values?

  1. AK-Nearest Neighbors (KNN) imputation, as it considers the values of similar data points.
  2. BZero imputation, as it is straightforward and often used for numerical features.
  3. CMedian imputation, as it is less sensitive to outliers and skewed distributions.
  4. DMode imputation, as it represents the most frequent value and is suitable for skewed data.
Show answer & explanation

Correct answer: C. Median imputation, as it is less sensitive to outliers and skewed distributions.

Median imputation is robust to skewed distributions and outliers because it uses the central value, which is not heavily influenced by extreme values, thus preserving the shape of the distribution better than the mean.

Why the other options are wrong

  • A. Incorrect. While KNN imputation can be effective, it is more computationally intensive and might introduce noise if the 'similar' data points are not truly representative in a highly skewed distribution. Median is a simpler, more robust first choice here.
  • B. Incorrect. Zero imputation can introduce bias and is generally not appropriate for features where zero is not a meaningful absence of value, especially with skewed data.
  • D. Incorrect. Mode imputation is typically used for categorical data, not continuous numerical data like income.

Median Imputation

A data imputation technique where missing values in a feature are replaced by the median value of the non-missing entries in that feature.

  • Robust to outliers and skewed distributions.
  • Preserves the original distribution shape better than mean imputation for skewed data.
  • Suitable for numerical data.

Memory trick: Don't 'mean' to skew, 'median' is the clean way through.

More Exploratory Data Analysis questions