Professional Data EngineerDesigning data processing systemsMedium

A financial services company needs to process daily batch jobs that involve complex transformations and aggregations on terabytes of historical transaction data. The jobs have strict Service Level Agreements (SLAs) for completion time and require a fully managed, serverless solution to minimize operational overhead. The solution must also be cost-effective, paying only for the resources consumed during job execution. Which Google Cloud service is most suitable for this scenario?

  1. ACloud SQL.
  2. BCompute Engine.
  3. CDataflow.
  4. DDataproc.
Show answer & explanation

Correct answer: C. Dataflow.

Dataflow is a fully managed, serverless service for executing Apache Beam pipelines. It automatically scales resources and optimizes job execution, making it ideal for complex batch transformations on terabytes of data with strict SLAs. Its pay-as-you-go model ensures cost-effectiveness by only charging for the resources used during the job's runtime, minimizing operational overhead.

Why the other options are wrong

  • A. Cloud SQL is a relational database and not designed for large-scale batch data transformations and aggregations.
  • B. Compute Engine provides raw VMs, requiring significant manual management for scaling, fault tolerance, and job orchestration, which increases operational overhead.
  • D. Dataproc is a managed Apache Spark/Hadoop service, but it requires managing clusters and is not fully serverless, incurring costs for idle clusters.

Serverless Batch Processing

Executing large-scale batch data transformations and aggregations using a fully managed, auto-scaling, and pay-per-use serverless service.

  • Minimizes operational overhead.
  • Scales automatically based on workload.
  • Cost-effective with pay-as-you-go pricing.

Memory trick: Dataflow: Data flows, serverless, saving cash.

More Designing data processing systems questions