AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A global logistics company receives shipment tracking updates from thousands of partners in various formats (CSV, XML, JSON). These files are uploaded to an S3 bucket several times a day. The company needs to consolidate, normalize, and enrich this data with internal reference tables before storing it in a unified Parquet format in another S3 bucket for analytical purposes. The processing should be cost-effective and handle schema evolution. Which AWS service should be used for this transformation?

  1. AAmazon Athena
  2. BAWS Step Functions with AWS Lambda
  3. CAmazon Kinesis Data Analytics for Apache Flink
  4. DAWS Glue ETL
Show answer & explanation

Correct answer: D. AWS Glue ETL

AWS Glue ETL is a serverless data integration service that is highly suitable for this scenario. It can read data in various formats from S3, perform complex transformations (consolidation, normalization, enrichment with joins), handle schema evolution through its Data Catalog, and write the processed data to S3 in Parquet format cost-effectively, as it's a pay-as-you-go service for job execution.

Why the other options are wrong

  • A. Amazon Athena is a query service for S3, not a transformation service. While it can query the data, it's not designed to consolidate, normalize, and enrich data into a new, transformed dataset.
  • B. While AWS Step Functions can orchestrate workflows, and Lambda can perform transformations, building a robust, scalable, and schema-evolution-aware solution for large-scale data transformation (potentially petabytes) using only Lambda would be significantly more complex and potentially less efficient than using a dedicated ETL service like AWS Glue.
  • C. Kinesis Data Analytics for Apache Flink is for real-time streaming data, not for batch processing of files uploaded 'several times a day' to S3.

AWS Glue ETL

A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development.

  • Serverless and scalable for batch ETL jobs.
  • Supports various data formats (CSV, XML, JSON) and schema evolution.
  • Integrates with Glue Data Catalog for metadata management.
  • Cost-effective: pay-as-you-go for job execution time.

Memory trick: Glue is the 'Librarian' normalizing your diverse data books.

More Data Ingestion and Transformation questions