CompTIA Data+ (DA0-002)Data MiningHard
A data scientist is preparing a dataset for a machine learning model that predicts housing prices. The dataset includes a 'Square_Footage' column, which has a few extremely high values that are clearly data entry errors (e.g., 100,000 sq ft for a typical residential home). These outliers can skew the model's training. Removing these records entirely is not feasible due to the limited dataset size. Which data cleansing technique should the data scientist use to mitigate the impact of these extreme outliers without deleting the records?
- AWinsorization
- BLog transformation
- CMode imputation
- DMean imputation
Show answer & explanationAnswer & explanation
Correct answer: A. Winsorization
Winsorization is a technique that limits extreme values in a dataset to a specified percentile, replacing outliers with the nearest non-outlier value. This reduces the influence of outliers without removing the data points entirely, which is crucial given the limited dataset size.
Why the other options are wrong
- B. Log transformation can reduce the skewness caused by outliers but doesn't directly 'correct' erroneous extreme values like Winsorization does by capping them.
- C. Mode imputation replaces missing values with the mode, not handles outliers.
- D. Mean imputation replaces missing values with the mean, not handles outliers.
Winsorization
A statistical method of limiting extreme values in a dataset to reduce the effect of spurious outliers. Outliers are replaced with the nearest non-outlier value.
- Reduces the influence of outliers without deleting data points.
- Often applied by capping values at a certain percentile (e.g., 5th and 95th).
- Preserves the number of observations in the dataset.
Memory trick: Outliers need careful handling, not always removal.