Google Associate Cloud EngineerDeploying and implementing a cloud solutionMedium
A data engineering team needs to process large datasets (terabytes) for analytical purposes using Apache Spark. They require a fully managed, cost-effective solution that can automatically scale compute resources up and down based on demand. Which Google Cloud service should they choose?
- ACompute Engine instances with manual Spark installation
- BDataproc
- CDataflow
- DBigQuery
Show answer & explanationAnswer & explanation
Correct answer: B. Dataproc
Dataproc is a fully managed service for running Apache Spark, Hadoop, Flink, and Presto clusters. It offers automatic scaling, cost-effectiveness, and integrates well with other Google Cloud services, making it ideal for processing large datasets with Spark.
Why the other options are wrong
- A. Manual Spark installation on Compute Engine requires significant management overhead and doesn't offer automatic scaling or cost-effectiveness of a managed service.
- C. Dataflow is a fully managed service for stream and batch data processing based on Apache Beam, but the requirement specifically mentions Apache Spark.
- D. BigQuery is a serverless data warehouse for analytics, not a processing engine for running custom Spark jobs.
Dataproc
A fully managed, highly scalable service for running Apache Spark, Hadoop, Flink, and Presto clusters on Google Cloud. It simplifies big data processing with auto-scaling and cost-effectiveness.
- Fully managed Spark, Hadoop, Flink, Presto.
- Auto-scaling clusters.
- Cost-effective (per-second billing).
- Integrates with other GCP services.
Memory trick: Dataproc: Your managed Spark engine for big data's start.