Professional Data EngineerDesigning data processing systemsMedium

A pharmaceutical company is building a data pipeline to process genomic sequencing data. This data is extremely large (petabytes per study) and requires complex, computationally intensive transformations (e.g., alignment, variant calling) that can take days to complete. The processing jobs are not time-sensitive but must be fault-tolerant and scale elastically to handle varying workloads and data volumes across different research projects. Which Google Cloud service is most suitable for orchestrating and executing these batch processing jobs?

  1. ACloud Functions
  2. BDataflow
  3. CCloud Composer
  4. DCloud Pub/Sub
Show answer & explanation

Correct answer: B. Dataflow

Dataflow is a fully managed service for executing Apache Beam pipelines, which are well-suited for large-scale batch processing. It provides automatic scaling, fault tolerance, and can handle petabyte-scale data with complex transformations, making it ideal for computationally intensive genomic data processing that can run for long durations.

Why the other options are wrong

  • A. Cloud Functions are serverless functions designed for short-lived, event-driven tasks, not for long-running, petabyte-scale batch processing.
  • C. Cloud Composer (managed Apache Airflow) is an orchestration service for workflows, but Dataflow is the execution engine best suited for the actual data processing tasks, not Composer itself.
  • D. Cloud Pub/Sub is a messaging service for ingestion and decoupling, not an execution engine for complex data transformations.

Dataflow for Large-scale Batch Processing

A fully managed Google Cloud service for executing Apache Beam pipelines, enabling highly scalable, fault-tolerant, and cost-effective batch data processing for petabyte-scale datasets.

  • Serverless and auto-scaling.
  • Supports complex transformations (Apache Beam).
  • Built-in fault tolerance and exactly-once processing (for streaming).
  • Ideal for large-scale ETL, ELT, and data-intensive computations.

Memory trick: For a massive data factory, you need a 'Dataflow' river that can carry huge loads, process them, and handle any bumps along the way.

More Designing data processing systems questions