CompTIA Data+ (DA0-002)Data MiningMedium

A data scientist is analyzing a customer feedback dataset. They notice that the 'Sentiment Score' column, which should range from -1 (negative) to 1 (positive), contains some values like -5, 2, and 'N/A'. To prepare this data for a machine learning model, which data cleansing technique would be most appropriate for handling the 'N/A' values?

  1. ADeduplication
  2. BData standardization
  3. COutlier detection and removal
  4. DImputation
Show answer & explanation

Correct answer: D. Imputation

Imputation is the process of replacing missing data points (like 'N/A' in this case) with substituted values. This allows the dataset to remain complete for analysis and model training, rather than discarding rows with missing information.

Why the other options are wrong

  • A. Deduplication identifies and removes duplicate records, which is unrelated to missing values.
  • B. Data standardization scales or transforms data to a common range, not handling missing values.
  • C. Outlier detection addresses values like -5 and 2, but not 'N/A' which represents missing data.

Data Imputation

Data imputation is the process of replacing missing data with substituted values. The goal is to fill in gaps in a dataset to maintain data integrity and enable complete analysis.

  • Common methods include mean, median, mode, or predictive imputation.
  • Helps maintain the dataset size and statistical power.
  • Choice of method depends on the nature of missing data and variable distribution.

Memory trick: Missing? Impute it! Outlier? Fix it! Duplicate? Delete it!

More Data Mining questions