CompTIA Data+ (DA0-002)Data MiningMedium
A data quality specialist is analyzing a dataset of customer feedback where users are asked to rate their experience on a scale of 1 to 5. They discover that some entries contain values like '6', 'Excellent', or 'N/A'. To prepare this data for numerical analysis, these invalid entries must be addressed. Which data cleansing technique is most appropriate to ensure all values conform to the expected 1-5 numerical scale?
- AStandardization and Validation
- BMissing Value Imputation
- COutlier Treatment
- DDeduplication
Show answer & explanationAnswer & explanation
Correct answer: A. Standardization and Validation
Standardization ensures values conform to a common format or scale, while validation checks if data meets predefined rules or constraints. Handling '6' (out of range), 'Excellent' (wrong data type), and 'N/A' (non-numeric) all fall under these techniques to enforce the 1-5 numerical scale.
Why the other options are wrong
- B. Missing value imputation is for filling in genuinely absent data points, not for correcting existing but invalid entries like '6' or 'Excellent'.
- C. Outlier treatment typically deals with extreme but valid numerical values that lie far from the central tendency, not with values that are fundamentally outside the defined range or type.
- D. Deduplication focuses on identifying and removing duplicate records, not on correcting invalid data values within a field.
Data Standardization
The process of transforming data into a common format, scale, or structure to ensure consistency and comparability across different sources or within a single dataset.
- Essential for integrating data from multiple sources.
- Includes converting units, data types, and formatting conventions.
- Often combined with data validation to enforce business rules.
Memory trick: Clean data is like a polished gem, ready for its big show.