CompTIA Data+ (DA0-002)Data MiningMedium
A data scientist is preparing a dataset for a predictive model where the 'Age' column contains several entries like 'twenty-five', '30 years old', and '45'. Before converting these to a numerical format, which data cleansing step is most critical to ensure consistency and prevent errors during the conversion?
- AMissing value imputation
- BOutlier removal
- CStandardization
- DDeduplication
Show answer & explanationAnswer & explanation
Correct answer: C. Standardization
Standardization is crucial here as the 'Age' column has inconsistent formats (words, numbers with text). Before any numerical conversion, these varying representations need to be brought to a single, consistent format (e.g., pure numbers) to ensure accurate and error-free processing.
Why the other options are wrong
- A. Missing value imputation deals with absent data, not inconsistent data formats.
- B. Outlier removal addresses extreme values, not inconsistent formats.
- D. Deduplication removes duplicate records, not inconsistent data formats within a field.
Data Standardization
The process of transforming data into a common format, scale, or unit to ensure consistency and comparability across different data sources or fields.
- Addresses inconsistencies in data representation, units, or naming conventions.
- Often involves converting data types, units of measurement, or text cases.
- A crucial step before data integration, analysis, or model training.
Memory trick: Cleansing data is like cleaning a room: first, standardize the mess.