Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium
A data science team needs to provision a temporary, scalable cluster of virtual machines to run Apache Spark jobs for large-scale data processing. The cluster should be easy to set up and tear down, and automatically manage the underlying infrastructure. Which Google Cloud service should they use?
- ACloud SQL with a large instance size
- BCompute Engine with custom VM images and startup scripts
- CDataproc
- DGoogle Kubernetes Engine (GKE) with Spark on Kubernetes
Show answer & explanationAnswer & explanation
Correct answer: C. Dataproc
Dataproc is a fully managed service for running Apache Spark, Hadoop, and other big data frameworks. It simplifies cluster provisioning, management, and scaling, making it ideal for temporary, scalable clusters for data processing jobs without managing individual VMs.
Why the other options are wrong
- A. Cloud SQL is a relational database and not designed for running Apache Spark jobs for large-scale data processing.
- B. Compute Engine requires manual setup and management of Spark on VMs, which is not 'easy to set up and tear down' or 'automatically manage'.
- D. GKE can run Spark, but Dataproc is specifically optimized and managed for big data ecosystems like Spark, offering a simpler experience for this use case.
Dataproc
A fully managed, highly scalable service for running Apache Spark, Apache Hadoop, Apache Flink, and other open-source data processing frameworks.
- Simplifies deployment and management of big data clusters.
- Supports various open-source data processing tools.
- Pay-per-use pricing model, ideal for temporary clusters.
- Integrates with other Google Cloud services like Cloud Storage and BigQuery.
Memory trick: Dataproc: It's like a 'data factory' in the cloud, churning out insights with Spark and Hadoop.