Professional Cloud ArchitectAnalyze and optimize technical and business processesMedium

A global logistics company is building a new real-time tracking system for its fleet of delivery vehicles. The system needs to ingest millions of location updates per minute, process them to identify route deviations, and store them for historical analysis. Latency for processing route deviations must be under 100 milliseconds. Which set of Google Cloud services would be most appropriate for this architecture?

  1. ACloud IoT Core for ingestion, Compute Engine for processing, Cloud Spanner for storage.
  2. BCloud Pub/Sub for ingestion, Cloud Functions for processing, Cloud SQL for storage.
  3. CCloud Storage for ingestion, Dataflow for processing, BigQuery for storage.
  4. DCloud Pub/Sub for ingestion, Dataflow for real-time processing, BigQuery for storage.
Show answer & explanation

Correct answer: D. Cloud Pub/Sub for ingestion, Dataflow for real-time processing, BigQuery for storage.

Cloud Pub/Sub is ideal for ingesting high-volume, real-time data streams. Dataflow provides a fully managed service for real-time, low-latency stream processing. BigQuery is a highly scalable and cost-effective data warehouse for storing and analyzing large datasets, including historical analysis.

Why the other options are wrong

  • A. Cloud IoT Core is good for device connectivity but Pub/Sub is more general-purpose for high-volume ingestion. Compute Engine requires more management for stream processing than Dataflow. Cloud Spanner is excellent for global transactional databases but is overkill and more expensive for purely historical analytical storage compared to BigQuery.
  • B. Cloud Functions are generally better suited for event-driven, short-lived tasks rather than continuous, high-throughput stream processing. Cloud SQL might struggle with the sheer volume and velocity for historical analysis compared to BigQuery.
  • C. Cloud Storage is not suitable for high-throughput real-time ingestion with millisecond latency requirements.

Real-time Data Ingestion & Processing

Architectural patterns and services designed to capture, process, and analyze data streams with minimal latency.

  • Ingestion: High-throughput message queuing (e.g., Pub/Sub).
  • Processing: Stream analytics engine (e.g., Dataflow, Flink).
  • Storage: Scalable analytical database (e.g., BigQuery) or low-latency NoSQL (e.g., Bigtable).
  • Focus on low-latency, scalability, and fault tolerance.

Memory trick: Pub/Sub gets the data, Dataflow processes the stream, BigQuery stores the scene.

More Analyze and optimize technical and business processes questions