Professional Data EngineerDesigning data processing systemsEasy
A media company collects vast amounts of video metadata, user engagement logs, and content consumption patterns. They need to build a data pipeline to process this data for personalized content recommendations, audience segmentation, and content optimization. The data arrives continuously and requires transformations and aggregations before being stored in a data warehouse. The solution must be cost-effective, serverless, and provide a unified programming model for both batch and streaming data to simplify development and maintenance. Which Google Cloud service best fits the data processing requirements for transformations and aggregations?
- ABigQuery ML
- BDataproc
- CCloud Functions
- DDataflow
Show answer & explanationAnswer & explanation
Correct answer: D. Dataflow
Dataflow is a fully managed, serverless service for executing Apache Beam pipelines. It provides a unified programming model for both batch and streaming data, making it ideal for continuous transformations and aggregations of diverse data types. Its auto-scaling and serverless nature ensure cost-effectiveness and simplified operations.
Why the other options are wrong
- A. BigQuery ML is for creating and executing machine learning models directly within BigQuery, not for general data transformations and aggregations in a pipeline before warehousing.
- B. Dataproc is a managed Apache Hadoop and Spark service, which requires cluster management and is not serverless in the same way as Dataflow, nor does it inherently provide a unified batch/streaming model like Beam.
- C. Cloud Functions are suitable for event-driven, short-lived tasks, not for continuous, large-scale data stream processing and complex transformations.
Dataflow (Apache Beam)
Google Cloud Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, enabling unified programming for both batch and streaming data processing with auto-scaling and high performance.
- Unified programming model for batch and streaming.
- Serverless and fully managed service.
- Auto-scaling for dynamic workloads.
- Supports complex data transformations and aggregations.
Memory trick: Data Flow, One Code, Both Streams.