Professional Cloud ArchitectManage and provision solution infrastructureMedium
A data analytics company needs to process large datasets (terabytes to petabytes) for batch analytics jobs. These jobs are intermittent but require significant compute resources when they run. They want a serverless, cost-effective solution that can automatically scale up and down and integrate well with other Google Cloud data services. Which service is the most suitable for this purpose?
- ACloud Functions
- BCompute Engine with custom scripts
- CCloud Dataproc
- DDataflow
Show answer & explanationAnswer & explanation
Correct answer: D. Dataflow
Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, making it ideal for large-scale batch and stream data processing. It automatically scales resources up and down as needed, integrating seamlessly with other Google Cloud data services, and offers a cost-effective pay-per-use model, aligning perfectly with the requirements for intermittent, large-scale batch analytics.
Why the other options are wrong
- A. Cloud Functions are for event-driven, short-lived functions, not for processing terabytes to petabytes of data in batch jobs.
- B. Compute Engine requires manual management of VMs and scaling, which is not serverless or as cost-effective for intermittent jobs.
- C. Cloud Dataproc is a managed Apache Hadoop/Spark service, suitable for big data, but Dataflow offers a more serverless and often more cost-effective model for intermittent batch processing without cluster management overhead.
Dataflow for Batch Analytics
Google Cloud Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, suitable for large-scale batch and stream data processing.
- Serverless and autoscaling
- Unified model for batch and streaming (Apache Beam)
- Cost-effective for intermittent, large-scale jobs
Memory trick: Dataflow flows through your big data, serverlessly.