Professional Data EngineerDesigning data processing systemsMedium

A research institution needs to process large scientific datasets, often terabytes in size, containing complex, unstructured and semi-structured data (e.g., sensor readings, genomic sequences, research papers). The processing involves custom algorithms implemented in Python, and jobs can run for several hours. They need a cost-effective solution for batch processing that allows them to scale compute resources on demand without managing servers. Which Google Cloud service is most suitable?

  1. ACompute Engine with custom scripts
  2. BCloud Functions
  3. CDataflow
  4. DDataproc
Show answer & explanation

Correct answer: C. Dataflow

Dataflow, based on Apache Beam, is a fully managed, serverless service ideal for batch and stream processing. It automatically scales resources and handles complex transformations, making it cost-effective as you only pay for processing capacity used. It supports custom Python code, fitting the research institution's needs perfectly for terabyte-scale batch jobs without server management.

Why the other options are wrong

  • A. Compute Engine requires manual server management, which the requirement explicitly aims to avoid ('without managing servers').
  • B. Cloud Functions are for lightweight, event-driven tasks, not long-running, terabyte-scale batch processing.
  • D. Dataproc is a managed Apache Hadoop/Spark service, but Dataflow is often more cost-effective and truly serverless for large-scale batch transformations with custom code, as it doesn't require cluster management.

Dataflow for Serverless Batch Processing

Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, enabling scalable and cost-effective batch and stream data processing without infrastructure management.

  • Serverless: No infrastructure to manage.
  • Autoscaling: Dynamically adjusts resources.
  • Supports Python, Java, Go Beam SDKs.
  • Ideal for complex ETL, batch analytics, and stream processing.

Memory trick: Dataflow lets your data 'flow' without server 'woes'.

More Designing data processing systems questions