CompTIA Data+ (DA0-002)Data MiningMedium
A data scientist is analyzing a customer feedback dataset. They notice that the 'Sentiment Score' column, which should range from -1 (negative) to 1 (positive), contains some values like -5, 2, and 'N/A'. To prepare this data for a machine learning model, which data cleansing technique would be most appropriate for handling the 'N/A' values?
- ADeduplication
- BData standardization
- COutlier detection and removal
- DImputation
Show answer & explanationAnswer & explanation
Correct answer: D. Imputation
Imputation is the process of replacing missing data points (like 'N/A' in this case) with substituted values. This allows the dataset to remain complete for analysis and model training, rather than discarding rows with missing information.
Why the other options are wrong
- A. Deduplication identifies and removes duplicate records, which is unrelated to missing values.
- B. Data standardization scales or transforms data to a common range, not handling missing values.
- C. Outlier detection addresses values like -5 and 2, but not 'N/A' which represents missing data.
Data Imputation
Data imputation is the process of replacing missing data with substituted values. The goal is to fill in gaps in a dataset to maintain data integrity and enable complete analysis.
- Common methods include mean, median, mode, or predictive imputation.
- Helps maintain the dataset size and statistical power.
- Choice of method depends on the nature of missing data and variable distribution.
Memory trick: Missing? Impute it! Outlier? Fix it! Duplicate? Delete it!