Professional Cloud ArchitectManage and provision solution infrastructureHard
A global IoT company collects vast amounts of time-series data from millions of devices. They need to ingest this data in real-time, process it to extract insights, and store it for long-term analysis. The solution must be highly scalable, fault-tolerant, and handle high throughput. Which combination of Google Cloud services would be most appropriate for this data pipeline?
- ACloud Functions, Cloud Run, and Cloud Firestore
- BCloud Load Balancing, Compute Engine, and BigQuery
- CPub/Sub, Dataflow, and Bigtable
- DCloud Storage, Dataflow, and Cloud SQL
Show answer & explanationAnswer & explanation
Correct answer: C. Pub/Sub, Dataflow, and Bigtable
Pub/Sub is ideal for real-time ingestion of high-volume event streams like IoT data. Dataflow provides a serverless platform for real-time processing and transformation of streaming data. Bigtable is a petabyte-scale, low-latency NoSQL database well-suited for storing large amounts of time-series data for analytics. This combination addresses all requirements for real-time, scalable, and fault-tolerant IoT data processing.
Why the other options are wrong
- A. Cloud Functions/Run are for event-driven microservices, not typically for high-throughput stream processing. Cloud Firestore is a document database, not ideal for petabyte-scale time-series data.
- B. Cloud Load Balancing and Compute Engine are for web serving, not directly for data ingestion and processing. BigQuery is for data warehousing, not optimal for low-latency time-series storage from raw IoT streams.
- D. Cloud SQL is not designed for petabyte-scale time-series data storage or high-throughput real-time ingestion.
IoT Data Pipeline on GCP
A common architecture for processing real-time IoT data involves Pub/Sub for ingestion, Dataflow for stream processing, and Bigtable for time-series storage.
- Pub/Sub: Real-time, scalable messaging for ingestion
- Dataflow: Serverless, unified stream/batch processing
- Bigtable: Petabyte-scale, low-latency NoSQL for time-series
Memory trick: Pub/Sub brings it in, Dataflow processes, Bigtable stores it for IoT.