AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A data engineer is integrating a new data source into an existing data lake. The new source provides daily CSV files, each containing hundreds of millions of records. These files need to be converted to Parquet format, partitioned by 'ingestion_date' and 'source_id', and then registered in the AWS Glue Data Catalog. The process must be scheduled to run daily. Which AWS service provides a serverless and managed way to orchestrate this entire workflow, including data conversion, partitioning, and catalog updates?

  1. AAmazon EMR with Apache Airflow
  2. BAmazon Kinesis Data Firehose
  3. CAWS Glue Workflows
  4. DAWS Step Functions
Show answer & explanation

Correct answer: C. AWS Glue Workflows

AWS Glue Workflows are designed to orchestrate complex ETL jobs, including multiple AWS Glue jobs, crawlers, and other AWS services. They can manage the dependencies between steps, such as running a Glue ETL job for conversion and partitioning, followed by a Glue Crawler to update the Data Catalog, all on a daily schedule in a serverless manner.

Why the other options are wrong

  • A. Amazon EMR with Apache Airflow can orchestrate workflows, but it's not serverless and requires managing an EMR cluster and Airflow environment, which increases operational overhead compared to AWS Glue Workflows.
  • B. Kinesis Data Firehose is for streaming data ingestion to S3, not for orchestrating complex batch ETL workflows involving format conversion, partitioning, and catalog updates.
  • D. AWS Step Functions can orchestrate workflows, but AWS Glue Workflows are specifically tailored for Glue ETL jobs and crawlers, offering more seamless integration and features for data catalog updates.

AWS Glue Workflows

A serverless orchestration service within AWS Glue that allows you to create and manage complex ETL pipelines involving multiple Glue jobs, crawlers, and triggers.

  • Orchestrates a sequence of AWS Glue jobs and crawlers.
  • Manages dependencies between tasks.
  • Provides scheduling capabilities for automated execution.
  • Serverless and fully managed.

Memory trick: Glue Workflows orchestrate, from CSV to Parquet, then Catalog!

More Data Engineering questions