AWS Certified Data Engineer – AssociateData Ingestion and TransformationHard

A global e-commerce company wants to centralize customer clickstream data from various regional websites. The data arrives as JSON objects, approximately 1KB each, with a peak ingestion rate of 50,000 records per second. The company needs to transform this data by filtering out bot traffic and enriching it with customer demographic information from a DynamoDB table before storing it in an Amazon S3 data lake in Parquet format. The solution must provide near real-time processing with low latency. Which AWS service is most appropriate for the *transformation* of this data?

  1. AAWS Lambda
  2. BAmazon Kinesis Data Firehose
  3. CAmazon Kinesis Data Analytics for Apache Flink
  4. DAWS Glue Streaming ETL
Show answer & explanation

Correct answer: D. AWS Glue Streaming ETL

AWS Glue Streaming ETL is a serverless, Apache Spark-based service designed for continuous ETL jobs on streaming data. It can perform complex transformations like filtering and joining with external data sources (DynamoDB) in near real-time, and write the output to S3 in Parquet format, making it ideal for the specified transformation requirements.

Why the other options are wrong

  • A. Lambda can process streaming data, but for a continuous stream of 50,000 records/second with complex enrichment and state management, building and managing a Lambda-based solution would be significantly more complex and potentially less cost-effective than a dedicated streaming ETL service like Glue Streaming ETL.
  • B. Kinesis Data Firehose offers basic transformations (e.g., format conversion, simple Lambda-based record transformation) but is not designed for complex, stateful, or multi-source enrichment scenarios required here.
  • C. Kinesis Data Analytics for Apache Flink is also excellent for real-time processing, but AWS Glue Streaming ETL provides a fully managed, serverless Spark environment which can be more straightforward for ETL-style transformations and writing to S3 in Parquet, especially when integrating with the Glue Data Catalog.

AWS Glue Streaming ETL

A serverless data integration service that continuously transforms and cleanses streaming data using Apache Spark.

  • Serverless and scales automatically for streaming workloads.
  • Built on Apache Spark, enabling complex transformations.
  • Integrates with Glue Data Catalog for schema management.
  • Ideal for continuous ETL on data from Kinesis, Kafka, etc.

Memory trick: Glue Streaming ETL is the 'Factory' for your raw data streams.

More Data Ingestion and Transformation questions