AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisEasy
A data scientist is preparing a dataset for a machine learning model. They notice that the 'CustomerAge' column contains several negative values, which are logically impossible. Which of the following data cleaning techniques is most appropriate to address this issue?
- AOutlier capping using the 99th percentile.
- BRemoving rows where 'CustomerAge' is negative.
- CApplying a log transformation to the 'CustomerAge' column.
- DReplacing negative values with the column's mean.
Show answer & explanationAnswer & explanation
Correct answer: B. Removing rows where 'CustomerAge' is negative.
Negative age values are physically impossible and indicate corrupted data. Removing these specific rows is the most direct and accurate way to handle such a clear data quality issue without distorting the remaining valid data.
Why the other options are wrong
- A. Capping is for extreme but plausible outliers, not impossible values.
- C. Log transformation handles skewed data but does not correct impossible negative values.
- D. Replacing with the mean would introduce synthetic, incorrect data where impossible values exist.
Handling Impossible Values
Impossible values are data points that contradict known constraints or physical laws, such as negative age or a percentage greater than 100%.
- Indicate data corruption or entry errors.
- Should generally be removed or corrected to ensure data integrity.
- Different from outliers, which are extreme but plausible values.
Memory trick: Impossible values - just toss them out!