Google Associate Cloud EngineerDeploying and implementing a cloud solutionEasy
A data science team needs to provision a temporary, scalable cluster of virtual machines to run Apache Spark jobs for large-scale data processing. They want a fully managed service that simplifies cluster deployment, management, and scaling, allowing them to focus on their data processing tasks. Which Google Cloud service should they use?
- AGoogle Kubernetes Engine (GKE)
- BDataproc
- CCloud Functions
- DCompute Engine
Show answer & explanationAnswer & explanation
Correct answer: B. Dataproc
Dataproc is a fully managed service for running Apache Spark, Hadoop, Flink, and other open-source data processing frameworks, simplifying cluster deployment, scaling, and management.
Why the other options are wrong
- A. GKE can run Spark but requires managing Kubernetes, and Dataproc offers a more specialized and simplified experience for Spark/Hadoop clusters.
- C. Cloud Functions is a serverless function platform for event-driven microservices, not for running large-scale Spark jobs on a cluster.
- D. Compute Engine requires manual provisioning and management of individual VMs, which is not 'fully managed' for a Spark cluster.
Dataproc
A fully managed service for running Apache Spark, Hadoop, Flink, and other open-source data processing frameworks on Google Cloud.
- Simplifies cluster deployment, management, and scaling.
- Integrates with other Google Cloud services like Cloud Storage and BigQuery.
- Supports both batch processing and streaming workloads.
Memory trick: Dataproc for Spark, BigQuery for warehouse, Dataflow for streams.