AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A financial services company needs to process large volumes of historical transaction data stored in Amazon S3. The data is in CSV format, partitioned by date, and totals hundreds of terabytes. The company needs to perform complex aggregations, join data with external reference tables, and cleanse inconsistent records. The processing can run daily during off-peak hours, and cost efficiency is a major concern. Which AWS service is best suited for this transformation task?

  1. AAmazon EMR with Apache Spark
  2. BAmazon Kinesis Data Analytics for Apache Flink
  3. CAWS Glue ETL
  4. DAmazon Athena
Show answer & explanation

Correct answer: C. AWS Glue ETL

AWS Glue ETL is a serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. It is well-suited for large-scale batch processing, complex transformations, and joining data from S3, making it highly cost-effective as you only pay for the resources consumed during the execution of your ETL jobs.

Why the other options are wrong

  • A. Amazon EMR with Apache Spark can perform these transformations but requires cluster management and can be less cost-effective for daily batch jobs compared to the serverless nature of AWS Glue ETL, especially for 'cost efficiency' as a major concern.
  • B. Kinesis Data Analytics for Apache Flink is designed for real-time stream processing, not large-scale historical batch transformations.
  • D. Amazon Athena is a query service for S3 data; while it can perform aggregations, it's not designed for complex, multi-step ETL processes involving data cleansing and transformations typical of an ETL job. It's more for interactive querying.

AWS Glue ETL

A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development.

  • Serverless and scales automatically.
  • Supports various data sources and targets, including S3.
  • Ideal for batch ETL jobs, complex transformations, data cleansing, and schema inference.
  • Cost-effective: pay-as-you-go for job execution time.

Memory trick: Glue is the 'Swiss Army Knife' for S3 data transformations.

More Data Ingestion and Transformation questions