Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium

A data engineering team needs to process large datasets (terabytes) for analytical purposes. They require a fully managed, serverless service that can run Apache Spark and Hadoop jobs without requiring them to manage clusters or servers. Which Google Cloud service meets these requirements?

  1. ACloud Dataproc
  2. BCloud Dataflow
  3. CBigQuery
  4. DCloud Composer
Show answer & explanation

Correct answer: A. Cloud Dataproc

Cloud Dataproc is a fully managed service for running Apache Spark, Hadoop, Flink, and Presto clusters. It allows users to run big data workloads without the operational overhead of managing clusters, fitting the 'serverless' operational model for these specific technologies.

Why the other options are wrong

  • B. Cloud Dataflow is a fully managed service for executing Apache Beam pipelines, not directly for Spark/Hadoop jobs.
  • C. BigQuery is a serverless data warehouse for analytical queries, not a platform for running Spark/Hadoop jobs.
  • D. Cloud Composer is a managed Apache Airflow service for orchestrating workflows, not for executing big data processing jobs itself.

Google Cloud Dataproc

A fully managed, highly scalable service for running Apache Spark, Hadoop, Flink, and Presto clusters on Google Cloud.

  • Simplifies big data processing by abstracting cluster management.
  • Supports open-source big data tools.
  • Offers auto-scaling and cost-efficiency for ephemeral or long-running clusters.

Memory trick: To 'process' big 'data' with Spark/Hadoop, you need 'Dataproc' to make it rock!

More Deploying and implementing a cloud solution questions