Professional Data EngineerEnsuring solution qualityHard

A healthcare organization is building a data pipeline to process protected health information (PHI) from various sources into a BigQuery data warehouse. They need to ensure that PHI is pseudonymized before it lands in BigQuery for analytics purposes, to comply with HIPAA regulations. The solution must be scalable, automated, and allow for re-identification of data under strict access controls when necessary for specific use cases. Which Google Cloud service and approach should they use to achieve this data governance requirement?

  1. ADataflow for custom PII detection and redaction logic
  2. BBigQuery row-level security policies and data masking
  3. CCloud Data Loss Prevention (DLP) with de-identification and re-identification capabilities
  4. DCloud KMS with envelope encryption for PHI fields
Show answer & explanation

Correct answer: C. Cloud Data Loss Prevention (DLP) with de-identification and re-identification capabilities

Cloud Data Loss Prevention (DLP) is explicitly designed for detecting, classifying, and de-identifying sensitive data like PHI. Its de-identification methods, such as tokenization or format-preserving encryption, allow for pseudonymization and also support re-identification when authorized, meeting the requirements for HIPAA compliance and controlled access to original data.

Why the other options are wrong

  • A. While Dataflow can process data, implementing custom PII detection and redaction logic is complex, error-prone, and less robust than a specialized service like DLP.
  • B. BigQuery row-level security and data masking control access to existing data but don't perform the initial pseudonymization/de-identification of PHI before storage.
  • D. Cloud KMS encrypts fields but doesn't inherently pseudonymize data or offer simple re-identification for analytics, making it less suitable for this specific requirement.

Cloud DLP De-identification

A set of techniques within Cloud Data Loss Prevention (DLP) to transform sensitive data into a less sensitive format while preserving utility, with options for re-identification.

  • Includes tokenization, pseudonymization, format-preserving encryption.
  • Used to comply with data privacy regulations like HIPAA, GDPR.
  • Can be configured to allow controlled re-identification of data.

Memory trick: DLP Detects, De-identifies, Deciphers.

More Ensuring solution quality questions