Professional Data EngineerBuilding and operationalizing data processing systemsMedium

A data engineering team is building a complex data pipeline that involves ingesting data from on-premises databases, transforming it with custom Python scripts, and loading the results into BigQuery. The pipeline has several interdependent stages, including data validation, enrichment, and aggregation, all scheduled to run daily. They need a managed service that can orchestrate these tasks, handle dependencies, and provide robust scheduling and monitoring capabilities. Which Google Cloud service should they choose?

  1. ACloud Data Fusion
  2. BCloud Composer
  3. CCloud SQL
  4. DDataflow
Show answer & explanation

Correct answer: B. Cloud Composer

Cloud Composer, based on Apache Airflow, is designed for programmatic orchestration of complex workflows with dependencies, scheduling, and monitoring, making it ideal for managing multi-stage data pipelines.

Why the other options are wrong

  • A. Cloud Data Fusion is an ETL service for visual data integration, not primarily for orchestrating complex, code-driven, multi-stage pipelines with custom scripts and dependencies.
  • C. Cloud SQL is a relational database service and is not used for data pipeline orchestration.
  • D. Dataflow is a data processing service for batch and streaming data, not an orchestration tool for managing an entire pipeline of interdependent tasks.

Cloud Composer

A fully managed workflow orchestration service built on Apache Airflow, enabling users to author, schedule, and monitor pipelines represented as Directed Acyclic Graphs (DAGs).

  • Ideal for complex, multi-stage data pipelines.
  • Programmatic workflow definition using Python.
  • Provides robust scheduling, dependency management, and monitoring.

Memory trick: The composer arranges the entire data symphony.

More Building and operationalizing data processing systems questions