Professional Cloud ArchitectManage and provision solution infrastructureHard

A global IoT company collects vast amounts of time-series data from millions of devices. They need to ingest this data in real-time, process it to extract insights, and store it for long-term analysis. The solution must be highly scalable, fault-tolerant, and handle high throughput. Which combination of Google Cloud services would be most appropriate for this data pipeline?

  1. ACloud Functions, Cloud Run, and Cloud Firestore
  2. BCloud Load Balancing, Compute Engine, and BigQuery
  3. CPub/Sub, Dataflow, and Bigtable
  4. DCloud Storage, Dataflow, and Cloud SQL
Show answer & explanation

Correct answer: C. Pub/Sub, Dataflow, and Bigtable

Pub/Sub is ideal for real-time ingestion of high-volume event streams like IoT data. Dataflow provides a serverless platform for real-time processing and transformation of streaming data. Bigtable is a petabyte-scale, low-latency NoSQL database well-suited for storing large amounts of time-series data for analytics. This combination addresses all requirements for real-time, scalable, and fault-tolerant IoT data processing.

Why the other options are wrong

  • A. Cloud Functions/Run are for event-driven microservices, not typically for high-throughput stream processing. Cloud Firestore is a document database, not ideal for petabyte-scale time-series data.
  • B. Cloud Load Balancing and Compute Engine are for web serving, not directly for data ingestion and processing. BigQuery is for data warehousing, not optimal for low-latency time-series storage from raw IoT streams.
  • D. Cloud SQL is not designed for petabyte-scale time-series data storage or high-throughput real-time ingestion.

IoT Data Pipeline on GCP

A common architecture for processing real-time IoT data involves Pub/Sub for ingestion, Dataflow for stream processing, and Bigtable for time-series storage.

  • Pub/Sub: Real-time, scalable messaging for ingestion
  • Dataflow: Serverless, unified stream/batch processing
  • Bigtable: Petabyte-scale, low-latency NoSQL for time-series

Memory trick: Pub/Sub brings it in, Dataflow processes, Bigtable stores it for IoT.

More Manage and provision solution infrastructure questions