Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium

A data engineering team needs to run daily batch jobs that process large datasets (terabytes) stored in Cloud Storage. These jobs involve complex transformations using Apache Spark. They need a cost-effective solution that allows them to provision and deprovision clusters on demand, paying only for the time the clusters are active. Which Google Cloud service should they use?

  1. ABigQuery
  2. BDataflow
  3. CCompute Engine with custom Spark installation
  4. DDataproc
Show answer & explanation

Correct answer: D. Dataproc

Dataproc is a fully managed service for running Apache Spark and Hadoop clusters. It allows for on-demand provisioning and deprovisioning, enabling users to pay only for active cluster time, making it cost-effective for batch processing of large datasets.

Why the other options are wrong

  • A. BigQuery is a serverless data warehouse for analytics, not a processing engine for custom Spark transformations on raw data.
  • B. Dataflow is a serverless service for stream and batch processing, but it uses Apache Beam, not directly Apache Spark, which was specified in the requirement.
  • C. Compute Engine requires manual setup and management of Spark clusters, increasing operational overhead and not being as cost-effective for on-demand use.

Google Cloud Dataproc

A fully managed, highly scalable service for running Apache Spark, Hadoop, Flink, and Presto clusters on Google Cloud.

  • Supports on-demand cluster creation and deletion.
  • Integrates with Cloud Storage, BigQuery, and other GCP services.
  • Cost-effective for ephemeral batch processing workloads.

Memory trick: Dataproc 'processes' data like a 'pro' with Spark.

More Deploying and implementing a cloud solution questions