Professional Data EngineerBuilding and operationalizing data processing systemsHard

A data engineering team is building a complex data pipeline that involves ingesting data from various sources (CSV, JSON, databases), applying multiple transformations (filtering, joining, aggregation), and loading the results into BigQuery. The pipeline has dependencies between tasks, needs to be scheduled daily, and requires robust error handling and monitoring. They also want to reuse common transformation logic across different datasets. Which Google Cloud service should they use for orchestration?

  1. ADataflow
  2. BCloud Functions
  3. CCloud Composer
  4. DCloud Run
Show answer & explanation

Correct answer: C. Cloud Composer

Cloud Composer (managed Apache Airflow) is specifically designed for orchestrating complex, multi-step data pipelines with dependencies, scheduling, robust error handling, and extensibility, making it ideal for managing ETL workflows across diverse sources and transformations.

Why the other options are wrong

  • A. Dataflow is a data processing engine for executing Beam pipelines, not an orchestration service for managing dependencies between different types of tasks (ingestion, transformation, loading).
  • B. Cloud Functions are typically for single-purpose, event-driven functions, not complex, multi-step pipeline orchestration.
  • D. Cloud Run is for deploying stateless containers, suitable for individual microservices but not for orchestrating a full DAG of data tasks.

Google Cloud Composer

Google Cloud Composer is a fully managed workflow orchestration service built on Apache Airflow, enabling users to author, schedule, and monitor complex data pipelines as Directed Acyclic Graphs (DAGs).

  • Managed Apache Airflow service
  • Orchestrates complex multi-step workflows
  • Supports Python for defining DAGs
  • Integrated with other Google Cloud services

Memory trick: Composer: Conducts the data symphony.

More Building and operationalizing data processing systems questions