Professional Data EngineerDesigning data processing systemsHard

A global gaming company collects telemetry data from millions of active users. This data, which includes game events, player progress, and in-game purchases, needs to be ingested in real-time, processed to detect anomalies and personalize user experiences, and then stored for historical analysis. The system must be highly available, fault-tolerant, and scale dynamically to handle sudden spikes in user activity during new game releases. Which Google Cloud services should be used to design this hybrid real-time and batch data pipeline?

  1. ACloud Tasks, App Engine, and Memorystore.
  2. BCloud IoT Core, Cloud Functions, and Cloud Bigtable.
  3. CCloud Pub/Sub, Dataflow, BigQuery, and Cloud Storage.
  4. DCloud SQL, Compute Engine, and Cloud Spanner.
Show answer & explanation

Correct answer: C. Cloud Pub/Sub, Dataflow, BigQuery, and Cloud Storage.

This scenario requires a robust hybrid real-time and batch data pipeline. Cloud Pub/Sub is excellent for real-time ingestion of high-volume telemetry data. Dataflow, with its unified programming model, can perform both real-time stream processing (for anomaly detection and personalization) and batch processing (for historical data backfills or complex aggregations). BigQuery serves as a scalable data warehouse for historical analysis, while Cloud Storage can store raw or semi-processed data for long-term retention or as a staging area for batch jobs.

Why the other options are wrong

  • A. Cloud Tasks is for asynchronous task execution, not real-time stream ingestion. App Engine is for web applications, not data pipelines. Memorystore is for caching, not persistent storage of historical data.
  • B. Cloud IoT Core is for device management, not general event ingestion. Cloud Functions are for short-lived, event-driven tasks, not continuous stream processing. Cloud Bigtable is for operational analytics, but BigQuery is better for general historical analysis.
  • D. Cloud SQL and Cloud Spanner are transactional databases, not ideal for high-volume telemetry ingestion and petabyte-scale analytical storage. Compute Engine requires significant management.

Hybrid Real-time & Batch Pipeline

A data pipeline combining real-time stream processing for immediate insights with batch processing for comprehensive historical analysis, designed for high availability and scalability.

  • Uses Pub/Sub for real-time ingestion.
  • Leverages Dataflow for unified stream and batch processing.
  • Stores data in BigQuery for analytics and Cloud Storage for raw/long-term.

Memory trick: Pub/Sub, Dataflow, BigQuery, Storage: The Hybrid's Core for Gaming Data.

More Designing data processing systems questions