Google Associate Cloud EngineerEnsuring successful operation of a cloud solutionHard

A data analytics team needs to process large datasets that are stored in Cloud Storage. The processing requires significant computational power but is intermittent and can tolerate some startup latency. The team wants to use a managed service that automatically provisions and scales resources as needed, and they prefer to pay only for the compute time consumed. Which compute option should they use?

  1. ACompute Engine with preemptible VMs
  2. BGoogle Kubernetes Engine (GKE) Standard
  3. CDataflow
  4. DCloud Functions
Show answer & explanation

Correct answer: C. Dataflow

Dataflow is a fully managed service for executing Apache Beam pipelines that automatically provisions and scales resources for batch and stream processing. It's ideal for large-scale data processing that can tolerate some startup latency and offers a pay-per-use model.

Why the other options are wrong

  • A. Compute Engine with preemptible VMs reduces cost but requires manual management of instances and scaling, which goes against the 'managed service' and 'automatically scales' requirements.
  • B. GKE Standard provides a managed Kubernetes environment, but managing the cluster and scaling worker nodes for intermittent batch jobs still involves more operational overhead than a fully managed data processing service.
  • D. Cloud Functions are designed for short-lived, event-driven functions and are not suitable for large-scale data processing jobs that require significant computational power over extended periods.

Cloud Dataflow

A fully managed service for executing Apache Beam pipelines, enabling scalable and cost-effective batch and stream data processing.

  • Fully managed and serverless
  • Auto-scaling for batch and stream processing
  • Pay-per-use, integrates with Cloud Storage

Memory trick: Dataflow flows through your big data, effortlessly scaling.

More Ensuring successful operation of a cloud solution questions