CompTIA Data+ (DA0-002)Data MiningHard
A data analyst is performing an exploratory data analysis on a sales dataset. They notice that the 'Revenue' column contains a few values that are significantly higher than the vast majority of other entries, potentially skewing statistical measures like the mean and standard deviation. These extreme values, while possibly legitimate, are distorting the overall distribution. Which data cleansing technique is most appropriate to mitigate the impact of these extreme values on statistical analysis without necessarily deleting them?
- AData normalization or winsorization
- BSchema validation
- CDeduplication
- DMissing value imputation
Show answer & explanationAnswer & explanation
Correct answer: A. Data normalization or winsorization
Data normalization (like log transformation) can reduce the impact of skewed distributions and outliers by transforming the data scale. Winsorization, specifically, caps extreme values at a certain percentile, effectively reducing their influence without removal. Both are appropriate for mitigating the impact of outliers while retaining data.
Why the other options are wrong
- B. Schema validation checks if data conforms to defined types and structures, not if individual numerical values are statistically extreme.
- C. Deduplication removes duplicate records, not extreme numerical values.
- D. Missing value imputation fills in absent data, not addresses existing extreme values.
Outlier Treatment (Winsorization/Normalization)
Techniques used to manage the impact of extreme data points (outliers) on statistical analysis, either by transforming their values (winsorization) or by changing the data's scale (normalization).
- Winsorization caps outliers at a specified percentile.
- Normalization (e.g., log transform) can reduce skewness caused by outliers.
- Aims to preserve data while reducing the distorting effect of extremes.
Memory trick: Outliers are odd, so either cut them, cap them, or change their view.