A logistics company needs to process real-time updates from thousands of delivery vehicles. Each vehicle sends location, speed, and status updates every few seconds. This data needs to be ingested, transformed, and then stored for both real-time operational dashboards and historical analysis. The company requires a fully managed, scalable solution that minimizes operational overhead and supports flexible schema evolution. Which set of Google Cloud services would be most appropriate for this data pipeline?
- ACloud Storage -> Dataproc -> BigQuery
- BCloud Pub/Sub -> Dataflow -> Cloud SQL
- CCloud Pub/Sub -> Dataflow -> Cloud Bigtable -> BigQuery
- DCloud Pub/Sub -> Cloud Functions -> Firestore
Show answer & explanationAnswer & explanation
Correct answer: C. Cloud Pub/Sub -> Dataflow -> Cloud Bigtable -> BigQuery
Cloud Pub/Sub provides scalable, real-time ingestion. Dataflow (Apache Beam) offers a fully managed service for stream processing, handling transformations and schema evolution. Cloud Bigtable is excellent for low-latency access to the latest vehicle data for operational dashboards, while BigQuery is perfect for cost-effective, petabyte-scale historical analysis.
Why the other options are wrong
- A. Cloud Storage and Dataproc are primarily for batch processing, not real-time stream ingestion and processing. This option lacks a real-time operational store.
- B. Cloud SQL is not suited for the high-throughput, low-latency operational data needs at scale, nor for petabyte historical analysis.
- D. Cloud Functions might be too limited for complex, high-volume stream transformations, and Firestore may not scale as cost-effectively as Bigtable/BigQuery for petabytes of time-series data.
Real-time to Historical Data Pipeline
This pipeline leverages Cloud Pub/Sub for ingestion, Dataflow for stream processing, Cloud Bigtable for real-time operational access, and BigQuery for historical analytical storage, providing a comprehensive solution for diverse data needs.
- Pub/Sub: Scalable message ingestion.
- Dataflow: Managed stream processing (ETL).
- Cloud Bigtable: Low-latency operational store.
- BigQuery: Petabyte-scale historical analytics.
Memory trick: Pub/Sub gets it IN, Dataflow makes it SHINE, Bigtable shows it NOW, BigQuery saves it FOREVER.