Professional Data EngineerDesigning data processing systemsHard

A global gaming company collects telemetry data from millions of active users. This data, consisting of events like game progress, purchases, and errors, arrives continuously at extremely high volumes (billions of events per day). The company needs to perform real-time aggregations for leaderboards and fraud detection, and also store the raw data for historical analysis and machine learning model training. The solution must be highly available and cost-effective. Which architecture should the company implement?

  1. ACloud Pub/Sub -> Cloud Functions -> Firestore for real-time, Cloud Storage for raw.
  2. BCloud Pub/Sub -> Dataflow (streaming) -> Cloud Bigtable for real-time aggregations, Dataflow (batch) to BigQuery for historical/ML.
  3. CCloud Pub/Sub -> Dataflow (streaming) -> Cloud SQL for real-time, Cloud Storage for raw.
  4. DCloud Storage (batch uploads) -> Dataproc -> BigQuery for both real-time and historical.
Show answer & explanation

Correct answer: B. Cloud Pub/Sub -> Dataflow (streaming) -> Cloud Bigtable for real-time aggregations, Dataflow (batch) to BigQuery for historical/ML.

This architecture is robust and scalable. Cloud Pub/Sub handles massive ingestion. Dataflow in streaming mode performs real-time aggregations and updates a low-latency store like Cloud Bigtable for leaderboards/fraud. The same Dataflow pipeline can write raw data to Cloud Storage, and then a Dataflow batch job or direct BigQuery ingestion can move this raw data to BigQuery for historical analysis and ML training, ensuring cost-effectiveness and scalability for both real-time and batch workloads.

Why the other options are wrong

  • A. Cloud Functions and Firestore might struggle with 'billions of events per day' for complex real-time aggregations and petabyte-scale storage, and Firestore is not ideal for ML training on raw data.
  • C. Cloud SQL is not suitable for billions of real-time events and aggregations due to its relational nature and potential scaling limitations for this specific use case.
  • D. Using Cloud Storage for batch uploads and Dataproc lacks the real-time processing capabilities required for leaderboards and fraud detection. Dataproc is primarily for batch processing, not streaming.

Hybrid Real-time & Batch Data Pipeline

A robust data pipeline combining streaming ingestion and processing for real-time insights with batch processing for historical analysis and machine learning, leveraging services like Pub/Sub, Dataflow, Bigtable, and BigQuery.

  • Pub/Sub for high-volume, real-time ingestion.
  • Dataflow for flexible, scalable streaming and batch processing.
  • Cloud Bigtable for low-latency operational data (aggregations).
  • BigQuery for petabyte-scale historical data and ML training.

Memory trick: Pub/Sub streams, Dataflow transforms, Bigtable shows NOW, BigQuery learns LATER.

More Designing data processing systems questions