CompTIA Data+ (DA0-002)Data Governance, Quality and ControlsHard

A data quality team runs a deduplication analysis on a customer database containing 50,000 records and identifies 3,000 records as duplicates of existing customer entries. What is the uniqueness rate of the dataset?

  1. A94%
  2. B85%
  3. C6%
  4. D97%
Show answer & explanation

Correct answer: A. 94%

Uniqueness rate = (Total records - Duplicate records) / Total records = (50,000 - 3,000) / 50,000 = 47,000 / 50,000 = 0.94, or 94%.

Why the other options are wrong

  • B. 85% does not correspond to any correct calculation using these figures.
  • C. 6% is the duplicate rate (3,000/50,000), not the uniqueness rate.
  • D. This would result from an incorrect subtraction of only 1,500 duplicates instead of 3,000.

Uniqueness (Data Quality Dimension)

The degree to which each real-world entity is represented only once in a dataset, with no duplicate records.

  • Uniqueness rate = (Total records − Duplicates) / Total records
  • Duplicate rate = Duplicates / Total records
  • Poor uniqueness inflates counts and skews aggregate reporting

Memory trick: 'CACTUS' - Completeness, Accuracy, Consistency, Timeliness, Uniqueness, Validity Standards

More Data Governance, Quality and Controls questions