Professional Data EngineerManaging and securing dataHard

A gaming company is using Cloud Spanner for its global leaderboard, which stores player IDs and scores. Due to privacy regulations, player IDs must be pseudonymized to prevent direct identification while still allowing for internal analytics that require joining with other datasets. The pseudonymization must be reversible for specific authorized processes (e.g., customer support). How should you implement this privacy requirement?

  1. AImplement BigQuery row-level security to restrict access to raw player IDs based on user roles.
  2. BReplace player IDs with randomly generated UUIDs and store the mapping in a separate, highly secured database.
  3. CUse a one-way cryptographic hash function (e.g., SHA256) on player IDs before storing them in Spanner.
  4. DEncrypt player IDs using Cloud KMS with a symmetric encryption key and store the encrypted values in Spanner.
Show answer & explanation

Correct answer: D. Encrypt player IDs using Cloud KMS with a symmetric encryption key and store the encrypted values in Spanner.

Encrypting player IDs with a symmetric encryption key managed by Cloud KMS allows for reversible pseudonymization. Authorized processes can decrypt the IDs, while unauthorized access to the Spanner data will only reveal ciphertext, preventing direct identification. This meets the reversible pseudonymization requirement.

Why the other options are wrong

  • A. Row-level security restricts access to rows, not column-level pseudonymization of data within a column. Also, Spanner does not natively support BigQuery-style row-level security.
  • B. Replacing with UUIDs and storing a mapping is a form of tokenization, which is reversible, but managing a separate 'highly secured database' for mappings adds complexity and potentially new security surface area. Cloud KMS provides a more integrated and managed solution for key management.
  • C. One-way hash functions are irreversible, which violates the requirement for reversible pseudonymization for authorized processes.

Reversible Pseudonymization with KMS

Reversible pseudonymization involves replacing direct identifiers with pseudonyms in a way that allows the original identifiers to be retrieved by authorized parties using a controlled key, often managed by a service like Cloud KMS.

  • Protects privacy by obscuring direct identifiers.
  • Allows data to be re-identified under strict controls.
  • Cloud KMS manages the encryption/decryption keys securely.
  • Useful for internal analytics requiring joins while maintaining privacy.

Memory trick: KMS has the KEY, for pseudonymity, you see!

More Managing and securing data questions