CompTIA Data+ (DA0-002)Data Governance, Quality and ControlsHard
A data quality team runs a deduplication analysis on a customer database containing 50,000 records and identifies 3,000 records as duplicates of existing customer entries. What is the uniqueness rate of the dataset?
- A94%
- B85%
- C6%
- D97%
Show answer & explanationAnswer & explanation
Correct answer: A. 94%
Uniqueness rate = (Total records - Duplicate records) / Total records = (50,000 - 3,000) / 50,000 = 47,000 / 50,000 = 0.94, or 94%.
Why the other options are wrong
- B. 85% does not correspond to any correct calculation using these figures.
- C. 6% is the duplicate rate (3,000/50,000), not the uniqueness rate.
- D. This would result from an incorrect subtraction of only 1,500 duplicates instead of 3,000.
Uniqueness (Data Quality Dimension)
The degree to which each real-world entity is represented only once in a dataset, with no duplicate records.
- Uniqueness rate = (Total records − Duplicates) / Total records
- Duplicate rate = Duplicates / Total records
- Poor uniqueness inflates counts and skews aggregate reporting
Memory trick: 'CACTUS' - Completeness, Accuracy, Consistency, Timeliness, Uniqueness, Validity Standards