Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium

A data engineering team needs to process large datasets (terabytes) for analytical purposes using Apache Spark. They require a fully managed, cost-effective solution that can automatically scale compute resources up and down based on demand. Which Google Cloud service should they choose?

  1. ACompute Engine instances with manual Spark installation
  2. BDataproc
  3. CDataflow
  4. DBigQuery
Show answer & explanation

Correct answer: B. Dataproc

Dataproc is a fully managed service for running Apache Spark, Hadoop, Flink, and Presto clusters. It offers automatic scaling, cost-effectiveness, and integrates well with other Google Cloud services, making it ideal for processing large datasets with Spark.

Why the other options are wrong

  • A. Manual Spark installation on Compute Engine requires significant management overhead and doesn't offer automatic scaling or cost-effectiveness of a managed service.
  • C. Dataflow is a fully managed service for stream and batch data processing based on Apache Beam, but the requirement specifically mentions Apache Spark.
  • D. BigQuery is a serverless data warehouse for analytics, not a processing engine for running custom Spark jobs.

Dataproc

A fully managed, highly scalable service for running Apache Spark, Hadoop, Flink, and Presto clusters on Google Cloud. It simplifies big data processing with auto-scaling and cost-effectiveness.

  • Fully managed Spark, Hadoop, Flink, Presto.
  • Auto-scaling clusters.
  • Cost-effective (per-second billing).
  • Integrates with other GCP services.

Memory trick: Dataproc: Your managed Spark engine for big data's start.

More Deploying and implementing a cloud solution questions