Professional Data EngineerOperationalizing machine learning modelsEasy

A healthcare provider is developing a machine learning model to assist in disease diagnosis. Due to strict regulatory compliance (HIPAA, GDPR), they must ensure that no personally identifiable information (PII) or protected health information (PHI) is inadvertently used during model training or exposed during inference. They need a mechanism to identify and redact sensitive data within their datasets before it reaches the ML pipeline. Which Google Cloud service should they integrate into their data preparation workflow?

  1. ACloud Identity and Access Management (IAM)
  2. BCloud Audit Logs
  3. CCloud Key Management Service (KMS)
  4. DCloud Data Loss Prevention (DLP)
Show answer & explanation

Correct answer: D. Cloud Data Loss Prevention (DLP)

Cloud Data Loss Prevention (DLP) is specifically designed to discover, classify, and protect sensitive data, including PII and PHI, across various data sources, making it ideal for redacting sensitive information before ML training.

Why the other options are wrong

  • A. Cloud IAM controls who can access resources, but it doesn't identify or redact sensitive data within the data itself.
  • B. Cloud Audit Logs record administrative activities and data access, but they don't prevent sensitive data from entering the ML pipeline.
  • C. Cloud KMS manages encryption keys, providing data at-rest and in-transit encryption, but it does not identify or redact sensitive data content.

Cloud Data Loss Prevention (DLP)

A fully managed service for discovering, classifying, and protecting sensitive data at scale, including PII, PHI, and financial data.

  • Detects over 150 types of sensitive data.
  • Supports de-identification, redaction, tokenization.
  • Works across various Google Cloud data sources.

Memory trick: DLP prevents data from leaking out.

More Operationalizing machine learning models questions