Professional Cloud ArchitectManage and provision solution infrastructureMedium

A data analytics company needs to process large datasets (terabytes to petabytes) for batch analytics jobs. These jobs are intermittent but require significant compute resources when they run. They want a serverless, cost-effective solution that can automatically scale up and down and integrate well with other Google Cloud data services. Which service is the most suitable for this purpose?

  1. ACloud Functions
  2. BCompute Engine with custom scripts
  3. CCloud Dataproc
  4. DDataflow
Show answer & explanation

Correct answer: D. Dataflow

Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, making it ideal for large-scale batch and stream data processing. It automatically scales resources up and down as needed, integrating seamlessly with other Google Cloud data services, and offers a cost-effective pay-per-use model, aligning perfectly with the requirements for intermittent, large-scale batch analytics.

Why the other options are wrong

  • A. Cloud Functions are for event-driven, short-lived functions, not for processing terabytes to petabytes of data in batch jobs.
  • B. Compute Engine requires manual management of VMs and scaling, which is not serverless or as cost-effective for intermittent jobs.
  • C. Cloud Dataproc is a managed Apache Hadoop/Spark service, suitable for big data, but Dataflow offers a more serverless and often more cost-effective model for intermittent batch processing without cluster management overhead.

Dataflow for Batch Analytics

Google Cloud Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, suitable for large-scale batch and stream data processing.

  • Serverless and autoscaling
  • Unified model for batch and streaming (Apache Beam)
  • Cost-effective for intermittent, large-scale jobs

Memory trick: Dataflow flows through your big data, serverlessly.

More Manage and provision solution infrastructure questions