AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy

A media company needs to process large volumes of video metadata (XML files) generated hourly. These files are stored in an Amazon S3 bucket and require parsing, validation against a schema, and transformation into a normalized JSON format before being loaded into a data warehouse for analytics. The solution must be cost-effective, serverless, and handle varying volumes of data efficiently. Which combination of AWS services is MOST suitable for this transformation?

  1. AAWS Glue ETL
  2. BAmazon EMR with Apache Spark
  3. CAmazon Kinesis Data Analytics for Apache Flink
  4. DAWS Lambda with custom Python scripts
Show answer & explanation

Correct answer: A. AWS Glue ETL

AWS Glue ETL is a serverless data integration service that is well-suited for batch processing of structured and semi-structured data like XML or JSON. It can automatically discover schema, parse data, perform transformations (validation, normalization), and write to a data warehouse or S3, meeting the cost-effective, serverless, and efficient processing requirements.

Why the other options are wrong

  • B. Amazon EMR with Apache Spark is powerful but requires cluster management, which might not be as cost-effective or serverless as Glue for hourly batch processing.
  • C. Amazon Kinesis Data Analytics for Apache Flink is for real-time streaming analytics, not for hourly batch processing of files stored in S3.
  • D. AWS Lambda can execute custom scripts, but managing the entire ETL pipeline, including schema discovery, complex transformations, and error handling for large XML files, would require significant custom development and orchestration, which Glue handles natively.

AWS Glue ETL

A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development.

  • Serverless and scalable
  • Supports various data formats (JSON, XML, CSV, Parquet, etc.)
  • Automated schema discovery with Glue Data Catalog
  • Apache Spark-based for powerful transformations

Memory trick: Glue ETL cleans and shapes your files like a skilled craftsman.

More Data Ingestion and Transformation questions