AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard

A data scientist is analyzing a large dataset of customer behavior, including features like 'browsing_duration', 'number_of_clicks', and 'purchase_amount'. During data cleaning, they identify a small number of records where 'browsing_duration' is recorded as 0, but 'number_of_clicks' is very high (e.g., 500+). These records contradict logical expectations (zero duration implies no activity, thus zero clicks). Removing these records would lead to significant data loss for other valuable features. Which data cleaning strategy is most appropriate for these contradictory records?

  1. ARemove both 'browsing_duration' and 'number_of_clicks' features from the dataset entirely.
  2. BReplace 'browsing_duration' with a small non-zero value (e.g., 1 second) and flag the record as potentially erroneous.
  3. CImpute 'browsing_duration' based on 'number_of_clicks' using a predictive model for consistent values.
  4. DReplace 'browsing_duration' with the mean of the feature, and keep 'number_of_clicks' as is.
Show answer & explanation

Correct answer: C. Impute 'browsing_duration' based on 'number_of_clicks' using a predictive model for consistent values.

Contradictory records where one feature (browsing_duration) logically depends on another (number_of_clicks) require a more sophisticated approach. Imputing 'browsing_duration' based on 'number_of_clicks' using a predictive model (e.g., regression) ensures consistency between these related features without losing the entire record, which is crucial if other features in the record are valuable.

Why the other options are wrong

  • A. Incorrect. Removing entire features due to a few contradictory records is a drastic measure that would lead to significant information loss for the entire dataset.
  • B. Incorrect. Replacing with a small non-zero value is an ad-hoc fix that doesn't resolve the underlying logical contradiction or ensure consistency with the high click count. Flagging is good, but doesn't fix the data.
  • D. Incorrect. Replacing 'browsing_duration' with the mean would not address the logical inconsistency with 'number_of_clicks' and might introduce further bias.

Logical Inconsistency Resolution

A data cleaning strategy to address records where values across multiple features contradict logical or domain-specific rules, often by using predictive imputation or rule-based correction to ensure internal consistency.

  • Involves multiple features with conflicting values.
  • Aims to preserve data while correcting logical errors.
  • Often requires domain expertise and predictive modeling for imputation.

Memory trick: Contradictory data's a tough call, predictive imputation saves it all.

More Exploratory Data Analysis questions