Professional Data EngineerDesigning data processing systemsHard

A pharmaceutical company needs to build a data pipeline to process genomic sequencing data. This data arrives in large batches (terabytes per day) and requires complex transformations, including alignment, variant calling, and annotation, before it can be used for research. The processing must be highly parallelizable, fault-tolerant, and able to scale dynamically based on the daily data volume, without requiring the team to manage underlying infrastructure. Which Google Cloud service is best suited for building this batch processing pipeline?

  1. ACloud Composer
  2. BCloud Functions
  3. CCloud Dataproc
  4. DCloud Dataflow
Show answer & explanation

Correct answer: D. Cloud Dataflow

Cloud Dataflow, based on Apache Beam, is a fully managed, serverless service designed for large-scale data processing, both batch and stream. It provides strong capabilities for complex transformations, dynamic scaling, and fault tolerance, making it ideal for terabyte-scale genomic data processing without infrastructure management.

Why the other options are wrong

  • A. Cloud Composer (Apache Airflow) is an orchestration service for managing workflows, not a data processing engine itself.
  • B. Cloud Functions are suitable for small, event-driven tasks, not for terabyte-scale, complex batch data transformations.
  • C. Cloud Dataproc is a managed Apache Hadoop and Spark service, requiring some cluster management and not fully serverless like Dataflow.

Dataflow for Serverless Batch Processing

Cloud Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, enabling large-scale, fault-tolerant batch and stream data processing with dynamic scaling.

  • Serverless and fully managed.
  • Supports both batch and stream processing.
  • Dynamically scales resources based on workload.

Memory trick: Dataflow makes data 'flow' smoothly, even in huge batches, without server 'fuss'.

More Designing data processing systems questions