CompTIA Data+ (DA0-002)Data MiningMedium

A data analyst is working with a large dataset of customer transactions. They notice that some records have missing values in the 'Purchase_Amount' column, which is critical for their analysis. Simply removing these records would lead to significant data loss. Which of the following data cleansing techniques would be most appropriate to address these missing values without discarding entire records?

  1. ANormalization
  2. BOutlier removal
  3. CImputation
  4. DDeduplication
Show answer & explanation

Correct answer: C. Imputation

Imputation is the process of replacing missing data with substituted values, allowing the retention of records that would otherwise be discarded due to missing information.

Why the other options are wrong

  • A. Normalization scales numerical data to a standard range, not addresses missing values.
  • B. Outlier removal deals with extreme values, not missing data points.
  • D. Deduplication focuses on identifying and removing duplicate records, not filling missing values.

Data Imputation

The process of replacing missing data with substituted values to maintain the integrity and size of a dataset.

  • Prevents data loss from records with missing values.
  • Common methods include mean, median, mode, or predictive imputation.
  • Can introduce bias if not carefully applied.

Memory trick: Missing pieces need careful filling.

More Data Mining questions