AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A data engineering team is building a machine learning pipeline that processes highly sensitive customer data. Before training models, PII (Personally Identifiable Information) must be masked or anonymized in the dataset. The data is stored in Amazon S3 as CSV files and needs to be processed in a serverless, scalable manner. The masking process involves replacing specific columns (e.g., 'customer_name', 'email_address') with fictitious but consistent values (e.g., 'customer_name_1', 'customer_name_2') across different rows for the same original PII, to maintain referential integrity for downstream analysis. Which approach is most suitable?

  1. AUse an AWS Glue ETL job with custom Python scripts to hash PII columns.
  2. BLeverage Amazon SageMaker Processing jobs with a PySpark script for data masking.
  3. CUtilize Amazon Comprehend's PII detection and redaction API for each record.
  4. DImplement AWS Lambda functions triggered by S3 events to process each CSV file line-by-line.
Show answer & explanation

Correct answer: A. Use an AWS Glue ETL job with custom Python scripts to hash PII columns.

AWS Glue ETL jobs, particularly with PySpark, are highly suitable for large-scale data transformation tasks like PII masking on data stored in S3. Custom Python scripts within Glue can implement sophisticated masking logic, such as consistent hashing or tokenization, to replace PII with fictitious but referentially consistent values. Glue is serverless and scalable, meeting the requirements.

Why the other options are wrong

  • B. Amazon SageMaker Processing jobs can be used, but AWS Glue is generally a more specialized and cost-effective service for serverless ETL and data preparation tasks outside of model training workflows, especially when the primary goal is data transformation for a data lake.
  • C. Amazon Comprehend's PII detection and redaction is primarily for identifying and redacting unstructured text. While it can detect PII, implementing consistent fictitious values across structured data for referential integrity is not its primary use case and would be inefficient for large CSV files.
  • D. Processing large CSV files line-by-line with AWS Lambda triggered by S3 events can be inefficient and complex to manage for large datasets, especially when maintaining referential integrity across multiple records. Lambda is better suited for smaller, event-driven tasks.

PII Masking with AWS Glue

Using AWS Glue ETL jobs to transform sensitive PII data in datasets by replacing it with anonymized, pseudonymized, or fictitious values, often maintaining referential integrity.

  • AWS Glue provides a serverless Spark-based environment.
  • Custom Python/PySpark scripts enable flexible masking logic.
  • Suitable for large-scale structured data in S3.
  • Can implement consistent hashing or tokenization for referential integrity.

Memory trick: Glue's custom scripts make PII disappear, leaving consistent stand-ins clear.

More Data Engineering questions