Professional Cloud ArchitectDesign and plan a cloud solution architectureEasy

A multinational corporation is migrating its legacy data analytics platform to Google Cloud. The platform processes petabytes of historical data daily, generating reports and insights for business intelligence. The workloads are primarily batch-oriented, highly parallelizable, and require a scalable, cost-effective solution with strong integration with other Google Cloud data services. The company prefers a serverless approach to minimize operational overhead. Which Google Cloud service should you recommend for processing these large datasets?

  1. ACompute Engine with custom Apache Spark cluster
  2. BCloud Functions
  3. CDataflow
  4. DDataproc
Show answer & explanation

Correct answer: C. Dataflow

Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, making it ideal for large-scale batch and streaming data processing, offering automatic scaling and cost-efficiency for petabyte-scale workloads.

Why the other options are wrong

  • A. Compute Engine with a custom Spark cluster requires significant operational overhead for cluster management, which goes against the 'serverless approach' requirement.
  • B. Cloud Functions are designed for event-driven, short-lived functions, not for processing petabytes of data in batch pipelines.
  • D. Dataproc is a managed Apache Spark and Hadoop service. While it supports large data processing, it's not fully serverless and requires some cluster management compared to Dataflow's purely serverless model.

Cloud Dataflow

A fully managed, serverless service on Google Cloud for executing Apache Beam pipelines, enabling scalable and cost-effective processing of large datasets in both batch and streaming modes.

  • Serverless and fully managed, minimizing operational overhead.
  • Supports Apache Beam, allowing unified batch and streaming processing.
  • Automatically scales resources based on workload demands.
  • Cost-effective for petabyte-scale data processing.

Memory trick: Dataflow streams through big data, serverless and fast.

More Design and plan a cloud solution architecture questions