CompTIA Data+ (DA0-002)Data Governance, Quality and ControlsMedium

A university research team publishes a dataset of survey responses for public use. Before release, the team strips all direct identifiers, aggregates responses into broad age and income brackets, and removes any combination of fields that could indirectly re-identify a participant, such that even the research team itself cannot trace a response back to an individual. Which technique is being applied?

  1. AAnonymization
  2. BTokenization
  3. CPseudonymization
  4. DDynamic data masking
Show answer & explanation

Correct answer: A. Anonymization

Anonymization irreversibly removes or alters identifying information so that individuals cannot be re-identified, even by the organization that processed the data. This differs from pseudonymization, where a reversible mapping (e.g., a key) could still allow re-identification.

Why the other options are wrong

  • B. Tokenization substitutes sensitive values with tokens mapped in a vault, which is also reversible, unlike anonymization.
  • C. Pseudonymization is reversible with a separately held key, so the organization could still re-identify individuals.
  • D. Dynamic data masking hides values from certain users at query time but doesn't permanently alter or aggregate the underlying data.

Anonymization

The irreversible process of removing or altering personal identifiers so that individuals cannot be re-identified from the data, even by the data holder.

  • Irreversible, unlike pseudonymization or tokenization
  • Often involves aggregation or generalization of fields
  • Fully anonymized data typically falls outside privacy regulation scope

Memory trick: Mask hides, Pseudonym swaps (reversible), Anonymize erases forever, Token vaults, Aggregate blends

More Data Governance, Quality and Controls questions