Google Associate Cloud EngineerDeploying and implementing a cloud solutionEasy

A data science team needs to provision a temporary, scalable cluster of virtual machines to run Apache Spark jobs for large-scale data processing. They want a fully managed service that simplifies cluster deployment, management, and scaling, allowing them to focus on their data processing tasks. Which Google Cloud service should they use?

  1. AGoogle Kubernetes Engine (GKE)
  2. BDataproc
  3. CCloud Functions
  4. DCompute Engine
Show answer & explanation

Correct answer: B. Dataproc

Dataproc is a fully managed service for running Apache Spark, Hadoop, Flink, and other open-source data processing frameworks, simplifying cluster deployment, scaling, and management.

Why the other options are wrong

  • A. GKE can run Spark but requires managing Kubernetes, and Dataproc offers a more specialized and simplified experience for Spark/Hadoop clusters.
  • C. Cloud Functions is a serverless function platform for event-driven microservices, not for running large-scale Spark jobs on a cluster.
  • D. Compute Engine requires manual provisioning and management of individual VMs, which is not 'fully managed' for a Spark cluster.

Dataproc

A fully managed service for running Apache Spark, Hadoop, Flink, and other open-source data processing frameworks on Google Cloud.

  • Simplifies cluster deployment, management, and scaling.
  • Integrates with other Google Cloud services like Cloud Storage and BigQuery.
  • Supports both batch processing and streaming workloads.

Memory trick: Dataproc for Spark, BigQuery for warehouse, Dataflow for streams.

More Deploying and implementing a cloud solution questions