AWS Certified Data Engineer – AssociateData Ingestion and TransformationHard

A global e-commerce company wants to centralize customer clickstream data from various regional websites into a single Amazon S3 data lake. The data arrives as high-volume, continuous streams of semi-structured JSON events. Before landing in S3, the company needs to enrich these events by joining them with customer demographic data from a DynamoDB table and filter out sensitive PII. The solution must be serverless, scalable, and minimize operational overhead. Which AWS service combination is best suited for this real-time ingestion and transformation pipeline?

  1. AAWS Glue streaming ETL job
  2. BAmazon Kinesis Data Streams with Amazon Kinesis Data Analytics for SQL
  3. CAmazon Kinesis Data Firehose with an AWS Lambda function for transformation
  4. DAmazon Managed Streaming for Apache Kafka (MSK) with AWS Step Functions
Show answer & explanation

Correct answer: A. AWS Glue streaming ETL job

AWS Glue streaming ETL jobs are designed for continuous, serverless processing of streaming data. They can read from Kinesis, perform complex transformations like joining with DynamoDB (using Glue connectors), filter data, and write to S3, all while managing schema evolution and operational aspects. This fits the requirement for high-volume, real-time, serverless transformation with enrichment and filtering.

Why the other options are wrong

  • B. Kinesis Data Analytics for SQL is good for simpler SQL-based aggregations and filtering, but complex joins with external NoSQL databases (DynamoDB) are challenging or impossible.
  • C. Kinesis Data Firehose with Lambda can perform transformations, but complex joins with external data sources like DynamoDB are more efficiently handled by a dedicated streaming ETL service like Glue.
  • D. MSK requires managing Kafka clusters, and Step Functions are for orchestrating workflows, not primarily for continuous streaming data transformation and enrichment itself.

AWS Glue Streaming ETL

A serverless data integration service that allows you to continuously process streaming data for ETL (Extract, Transform, Load) operations.

  • Processes data from Kinesis and Kafka.
  • Supports complex transformations, joins, and filtering.
  • Automatically scales and manages infrastructure.

Memory trick: Glue streams transform, enrich, and filter data on the fly.

More Data Ingestion and Transformation questions