CompTIA Data+ (DA0-002)Data MiningMedium
A data scientist is working with a dataset containing customer transaction IDs. They notice that the ID column, which should be unique for each transaction, contains several duplicate entries. If these duplicates are not removed, they could lead to incorrect aggregations and skewed analysis. What is the most direct and effective data cleansing technique to address this issue?
- ADeduplication
- BOutlier detection
- CMissing value imputation
- DData transformation
Show answer & explanationAnswer & explanation
Correct answer: A. Deduplication
Deduplication is the specific process of identifying and removing duplicate records from a dataset. In this scenario, where unique transaction IDs are expected, removing duplicate entries is the most direct and appropriate cleansing technique.
Why the other options are wrong
- B. Outlier detection focuses on extreme values, not identical records.
- C. Missing value imputation deals with absent data, not redundant existing data.
- D. Data transformation changes the format or structure of data, but doesn't inherently remove duplicate records unless specifically programmed to do so as part of a broader transformation rule.
Deduplication
A data cleansing process that identifies and removes duplicate records or entries from a dataset, ensuring uniqueness based on specific keys or attributes.
- Crucial for maintaining data integrity and accuracy.
- Can be performed based on exact matches or fuzzy matching algorithms.
- Prevents overcounting and skewed analysis in reporting and models.
Memory trick: Clean data, clear insights: no mess, no fuss, just truth.