CompTIA DataSys+ (DS0-001)Data and Database SecurityHard
A company is developing a new analytics platform that requires processing large volumes of sensitive customer data. To protect the privacy of individuals while still allowing data scientists to perform statistical analysis, the company wants to transform the data so that individual records cannot be linked back to a specific person, but the statistical properties of the dataset remain intact. Which data transformation technique would best achieve this goal?
- ATokenization
- BHashing
- CAnonymization
- DEncryption
Show answer & explanationAnswer & explanation
Correct answer: C. Anonymization
Anonymization is the process of removing or modifying personally identifiable information (PII) from a dataset so that individuals cannot be directly or indirectly identified. The key is that it preserves the utility of the data for statistical analysis while making re-identification practically impossible, which matches the requirement.
Why the other options are wrong
- A. Tokenization replaces sensitive data with non-sensitive substitutes, but often a link or vault exists to reverse it, or the tokens themselves might allow re-identification with other data.
- B. Hashing creates a one-way digest, useful for integrity checks or password storage, but doesn't retain statistical properties or allow analysis of the original data.
- D. Encryption renders data unreadable without a key, but if decrypted, it's still identifiable; it doesn't remove the link.
Anonymization
Anonymization is a data privacy technique that transforms data to prevent the re-identification of individuals while preserving the data's analytical utility. It typically involves removing or altering direct and indirect identifiers.
- Prevents re-identification of individuals.
- Retains statistical value of the data.
- Often considered irreversible.
- Techniques include generalization, suppression, perturbation.
Memory trick: Anonymization makes data nameless but useful.