Professional Data EngineerDesigning data processing systemsEasy

A media company collects vast amounts of video metadata, user engagement logs, and content consumption patterns. They need to build a data pipeline to process this data for personalized content recommendations, audience segmentation, and content optimization. The data arrives continuously and requires transformations and aggregations before being stored in a data warehouse. The solution must be cost-effective, serverless, and provide a unified programming model for both batch and streaming data to simplify development and maintenance. Which Google Cloud service best fits the data processing requirements for transformations and aggregations?

  1. ABigQuery ML
  2. BDataproc
  3. CCloud Functions
  4. DDataflow
Show answer & explanation

Correct answer: D. Dataflow

Dataflow is a fully managed, serverless service for executing Apache Beam pipelines. It provides a unified programming model for both batch and streaming data, making it ideal for continuous transformations and aggregations of diverse data types. Its auto-scaling and serverless nature ensure cost-effectiveness and simplified operations.

Why the other options are wrong

  • A. BigQuery ML is for creating and executing machine learning models directly within BigQuery, not for general data transformations and aggregations in a pipeline before warehousing.
  • B. Dataproc is a managed Apache Hadoop and Spark service, which requires cluster management and is not serverless in the same way as Dataflow, nor does it inherently provide a unified batch/streaming model like Beam.
  • C. Cloud Functions are suitable for event-driven, short-lived tasks, not for continuous, large-scale data stream processing and complex transformations.

Dataflow (Apache Beam)

Google Cloud Dataflow is a fully managed, serverless service for executing Apache Beam pipelines, enabling unified programming for both batch and streaming data processing with auto-scaling and high performance.

  • Unified programming model for batch and streaming.
  • Serverless and fully managed service.
  • Auto-scaling for dynamic workloads.
  • Supports complex data transformations and aggregations.

Memory trick: Data Flow, One Code, Both Streams.

More Designing data processing systems questions