Professional Data EngineerDesigning data processing systemsHard

A global ride-sharing company processes millions of GPS coordinates per second from active vehicles. This time-series data needs to be ingested continuously, stored efficiently for at least 30 days for immediate operational analysis (e.g., real-time traffic, dispatch optimization), and then archived for long-term historical analysis spanning several years. The system must support high write throughput and low-latency reads for recent data. Which combination of Google Cloud services should you recommend?

  1. ACloud Tasks for ingestion, Memorystore for recent data, and Dataproc for long-term archival.
  2. BCloud IoT Core for ingestion, Cloud SQL for recent data, and Cloud Storage for long-term archival.
  3. CCloud Functions for ingestion, Cloud Firestore for recent data, and Cloud Spanner for long-term archival.
  4. DCloud Pub/Sub for ingestion, Cloud Bigtable for recent data, and BigQuery for long-term archival.
Show answer & explanation

Correct answer: D. Cloud Pub/Sub for ingestion, Cloud Bigtable for recent data, and BigQuery for long-term archival.

This scenario describes a common pattern for time-series data: a hot path for immediate operational use and a cold path for long-term historical analysis. Cloud Pub/Sub is ideal for ingesting millions of messages per second. Cloud Bigtable is purpose-built for high-throughput, low-latency time-series data, making it perfect for storing 30 days of active GPS coordinates for operational analysis. BigQuery is then the most cost-effective and scalable solution for long-term archival and complex analytical queries over years of data.

Why the other options are wrong

  • A. Cloud Tasks is for asynchronous task execution. Memorystore is an in-memory cache, not persistent storage. Dataproc is for batch processing on clusters, not a long-term archival database.
  • B. Cloud IoT Core is for device management, not general Pub/Sub. Cloud SQL cannot handle millions of writes per second for time-series. Cloud Storage is for raw files, not a queryable data warehouse for long-term analysis like BigQuery.
  • C. Cloud Functions are for event-driven, short-lived tasks, not continuous ingestion. Cloud Firestore is a document database, not optimized for high-throughput time-series. Cloud Spanner is a transactional relational database, not suited for long-term, cost-effective analytical archival.

Time-Series Hot/Cold Path

An architecture pattern for time-series data where recent, frequently accessed data resides in a low-latency store (hot path) and older, less frequently accessed data is moved to a cost-effective archival store (cold path).

  • Hot path for immediate operational queries.
  • Cold path for long-term historical analysis.
  • Optimizes for both performance and cost.

Memory trick: Pub/Sub for Push, Bigtable for Hot, BigQuery for Cold Storage.

More Designing data processing systems questions