A global gaming company uses Cloud Spanner for its global leaderboard, which stores player IDs and high scores. To comply with privacy regulations, they need to ensure that actual player IDs are not stored directly in the leaderboard table but can be reversibly linked back to real player identities when necessary (e.g., for customer support or fraud investigation). The mapping between the pseudonymized ID and the real ID must be highly secure and centrally managed. Which approach should the data engineering team implement?
- AEncrypt real player IDs at the application layer before storing them in Spanner, managing keys in Cloud KMS.
- BUse Cloud Key Management Service (KMS) to perform deterministic encryption on player IDs for storage in Spanner.
- CCreate a separate, highly restricted Cloud SQL database to store the mapping between pseudonymized IDs and real IDs.
- DStore real player IDs in Spanner, but use BigQuery data masking to hide them from most users.
Show answer & explanationAnswer & explanation
Correct answer: B. Use Cloud Key Management Service (KMS) to perform deterministic encryption on player IDs for storage in Spanner.
Deterministic encryption using Cloud KMS is the most suitable approach for reversible pseudonymization. It allows the same input (real player ID) to always produce the same encrypted output (pseudonymized ID), enabling joins or lookups across tables in Spanner without exposing the original ID. The encryption is reversible for authorized users and managed securely by Cloud KMS.
Why the other options are wrong
- A. Application-layer encryption might not be deterministic, making it hard to join or look up pseudonymized IDs. While keys in KMS are good, the encryption method is key.
- C. Using a separate Cloud SQL database for mapping adds complexity, introduces another system to manage, and might have performance implications for lookups, whereas deterministic encryption keeps the pseudonymization within the data itself.
- D. BigQuery data masking is for BigQuery, not Cloud Spanner, and it doesn't pseudonymize the data at rest.
Reversible Pseudonymization with Cloud KMS (Deterministic Encryption)
Reversible pseudonymization using Cloud KMS involves encrypting sensitive identifiers (like PII) with a deterministic encryption algorithm. This means the same input always produces the same encrypted output, allowing for consistent lookups or joins on the pseudonymized data while enabling authorized decryption back to the original identifier when needed.
- Uses deterministic encryption for consistent pseudonymized values.
- Enables joins and analytics on pseudonymized data.
- Encryption keys are securely managed by Cloud KMS.
- Allows authorized reversal to the original data for specific purposes (e.g., support).
Memory trick: KMS gives a consistent mask, when you need to unmask, it's there.