CompTIA Data+ (DA0-002)Data MiningMedium

A data scientist is preparing a dataset for a predictive model where the 'Age' column contains several entries like 'twenty-five', '30 years old', and '45'. Before converting these to a numerical format, which data cleansing step is most critical to ensure consistency and prevent errors during the conversion?

  1. AMissing value imputation
  2. BOutlier removal
  3. CStandardization
  4. DDeduplication
Show answer & explanation

Correct answer: C. Standardization

Standardization is crucial here as the 'Age' column has inconsistent formats (words, numbers with text). Before any numerical conversion, these varying representations need to be brought to a single, consistent format (e.g., pure numbers) to ensure accurate and error-free processing.

Why the other options are wrong

  • A. Missing value imputation deals with absent data, not inconsistent data formats.
  • B. Outlier removal addresses extreme values, not inconsistent formats.
  • D. Deduplication removes duplicate records, not inconsistent data formats within a field.

Data Standardization

The process of transforming data into a common format, scale, or unit to ensure consistency and comparability across different data sources or fields.

  • Addresses inconsistencies in data representation, units, or naming conventions.
  • Often involves converting data types, units of measurement, or text cases.
  • A crucial step before data integration, analysis, or model training.

Memory trick: Cleansing data is like cleaning a room: first, standardize the mess.

More Data Mining questions