Certified Cloud Security Professional (CCSP)Cloud Data SecurityHard
A financial institution is implementing a cloud-based data lake to consolidate various data sources, including transactional data, customer demographics, and market feeds. Due to regulatory compliance requirements (e.g., GDPR, CCPA), the institution must ensure that individual customer data within the data lake cannot be directly linked back to a specific person, even if the data is compromised. However, the data must still be useful for statistical analysis and trend identification. Which data protection technique is most appropriate for this scenario?
- APseudonymization
- BEncryption
- CData Masking
- DTokenization
Show answer & explanationAnswer & explanation
Correct answer: A. Pseudonymization
Pseudonymization replaces personally identifiable information with artificial identifiers (pseudonyms) while maintaining the ability to re-identify the data if necessary, usually with a separate key or mechanism. This allows for statistical analysis without direct identification, meeting the requirement to prevent direct linking while allowing for utility.
Why the other options are wrong
- B. Encryption protects data confidentiality but doesn't inherently prevent re-identification if the key is compromised and the original data structure remains.
- C. Data masking replaces sensitive data with fictitious but realistic data, often for testing or training environments, making it difficult to reverse, but often less useful for statistical analysis of real trends than pseudonymized data.
- D. Tokenization replaces sensitive data with a non-sensitive equivalent (token), often used for payment card data, but typically aims for a one-to-one mapping that can still be de-tokenized, making re-identification possible with the token vault.
Pseudonymization
A data protection technique where personally identifiable information (PII) is replaced with an artificial identifier (pseudonym), making it difficult to attribute the data to a specific individual without additional information, while preserving the data's utility for analysis.
- Replaces PII with pseudonyms.
- Reversible with additional information (e.g., a key).
- Balances privacy with data utility for analytics.
Memory trick: Pseudonyms Preserve Privacy, Permit Patterns.