AWS Certified Machine Learning – SpecialtyData EngineeringHard

A healthcare organization is setting up a new data pipeline for machine learning, handling sensitive patient data. Due to HIPAA compliance requirements, all data must be pseudonymized before it is used for model training. This involves replacing direct identifiers (e.g., patient names, social security numbers) with artificial identifiers while maintaining referential integrity for analytical purposes. Which data transformation technique is most appropriate to meet this compliance requirement?

  1. AData imputation
  2. BData masking
  3. CData aggregation
  4. DFeature scaling
Show answer & explanation

Correct answer: B. Data masking

Data masking (specifically, pseudonymization) is the most appropriate technique. It replaces sensitive, direct identifiers with realistic but false data or artificial identifiers. This satisfies compliance requirements like HIPAA by protecting individual privacy while still allowing the data to be used for analysis and machine learning, often by maintaining the format and referential integrity of the original data.

Why the other options are wrong

  • A. Data imputation fills in missing values in a dataset. It is not a technique for anonymizing or pseudonymizing sensitive identifiers.
  • C. Data aggregation combines data to a higher level (e.g., counts, sums) and can reduce privacy risk, but it loses individual record detail and doesn't maintain referential integrity for specific patient analysis.
  • D. Feature scaling (e.g., normalization, standardization) transforms numerical features to a common range, which is for model performance, not for privacy compliance of identifiers.

Pseudonymization (Data Masking)

A data privacy technique where direct identifiers are replaced with artificial identifiers, ensuring individuals cannot be identified without additional information, while preserving data utility.

  • Key for compliance (e.g., HIPAA, GDPR).
  • Replaces sensitive data with non-sensitive substitutes.
  • Can maintain format and referential integrity.
  • Differs from anonymization (which aims for irreversible de-identification).
  • Often involves tokenization or cryptographic hashing for identifiers.

Memory trick: Masking hides the identities, keeping data safe and useful.

More Data Engineering questions