AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium

A data scientist is performing exploratory data analysis on a dataset containing customer transaction records. They observe that the 'TransactionAmount' column has a large number of missing values (approximately 30%). After inspecting the data, they find no discernible pattern or reason for these missing values; they appear to be randomly distributed. Which imputation strategy would be most suitable to handle these missing values while minimizing bias and preserving the variability of the original data?

  1. AMultiple imputation by chained equations (MICE)
  2. BDeletion of rows with missing values
  3. CMedian imputation
  4. DMean imputation
Show answer & explanation

Correct answer: A. Multiple imputation by chained equations (MICE)

With a high percentage of missing values (30%) and no discernible pattern (Missing At Random), simple imputation methods like mean/median imputation can severely underestimate variance and distort relationships. Deletion would lead to significant data loss. MICE is a sophisticated technique that imputes missing values multiple times, accounting for uncertainty and preserving data variability and relationships more effectively.

Why the other options are wrong

  • B. Deleting 30% of rows would result in a substantial loss of valuable data.
  • C. Median imputation also reduces variance and can introduce bias, similar to mean imputation.
  • D. Mean imputation reduces variance and can bias relationships, especially with 30% missing data.

Multiple Imputation by Chained Equations (MICE)

A sophisticated imputation technique that iteratively imputes missing data using a series of regression models, creating multiple complete datasets to account for imputation uncertainty.

  • Handles Missing At Random (MAR) and Missing Completely At Random (MCAR) data.
  • Preserves variability and relationships better than single imputation.
  • Generates multiple imputed datasets, then pools results.
  • Computationally more intensive but provides more accurate inferences.

Memory trick: MICE makes multiple models to make missing data meaningful.

More Exploratory Data Analysis questions