Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium
A data engineering team needs to run daily batch jobs that process large datasets (terabytes) stored in Cloud Storage. These jobs involve complex transformations using Apache Spark. They need a cost-effective solution that allows them to provision and deprovision clusters on demand, paying only for the time the clusters are active. Which Google Cloud service should they use?
- ABigQuery
- BDataflow
- CCompute Engine with custom Spark installation
- DDataproc
Show answer & explanationAnswer & explanation
Correct answer: D. Dataproc
Dataproc is a fully managed service for running Apache Spark and Hadoop clusters. It allows for on-demand provisioning and deprovisioning, enabling users to pay only for active cluster time, making it cost-effective for batch processing of large datasets.
Why the other options are wrong
- A. BigQuery is a serverless data warehouse for analytics, not a processing engine for custom Spark transformations on raw data.
- B. Dataflow is a serverless service for stream and batch processing, but it uses Apache Beam, not directly Apache Spark, which was specified in the requirement.
- C. Compute Engine requires manual setup and management of Spark clusters, increasing operational overhead and not being as cost-effective for on-demand use.
Google Cloud Dataproc
A fully managed, highly scalable service for running Apache Spark, Hadoop, Flink, and Presto clusters on Google Cloud.
- Supports on-demand cluster creation and deletion.
- Integrates with Cloud Storage, BigQuery, and other GCP services.
- Cost-effective for ephemeral batch processing workloads.
Memory trick: Dataproc 'processes' data like a 'pro' with Spark.