A global IoT company collects vast amounts of time-series data from millions of devices worldwide. They need a highly scalable, real-time data ingestion pipeline that can handle millions of events per second, process the data with custom business logic, and store it for both real-time dashboards and long-term analytics. Which combination of Google Cloud services should they use?
- ACloud Pub/Sub, Dataflow, BigQuery, Cloud Bigtable
- BCloud IoT Core, Cloud Run, Cloud SQL
- CCloud Pub/Sub, Cloud Functions, Cloud Storage
- DCloud Spanner, Cloud Dataflow, Looker Studio
Show answer & explanationAnswer & explanation
Correct answer: A. Cloud Pub/Sub, Dataflow, BigQuery, Cloud Bigtable
Cloud Pub/Sub provides scalable, real-time event ingestion. Dataflow can process streaming data with custom business logic and windowing. BigQuery is ideal for long-term analytics, and Cloud Bigtable is excellent for low-latency access for real-time dashboards with time-series data. This combination addresses all requirements for a global IoT data pipeline.
Why the other options are wrong
- B. Cloud IoT Core is for device connection and management. Cloud Run is for stateless containers, not typically for continuous stream processing like Dataflow. Cloud SQL is a relational database and not suitable for petabytes of time-series data with millions of events per second.
- C. Cloud Functions are typically for short-lived, single-purpose functions, which might not be ideal for continuous, complex stream processing. Cloud Storage is object storage, not optimized for real-time dashboard queries on time-series data.
- D. Cloud Spanner is a globally distributed relational database, not optimized for raw, high-volume time-series data ingestion and storage. Looker Studio is a visualization tool, not a processing or storage component for the pipeline itself.
IoT Data Pipeline (Pub/Sub, Dataflow, BigQuery, Bigtable)
A common Google Cloud architecture for processing IoT data involves Cloud Pub/Sub for scalable ingestion, Dataflow for real-time stream processing and transformation, BigQuery for long-term analytics and reporting, and Cloud Bigtable for low-latency access to time-series data for real-time dashboards.
- Cloud Pub/Sub: Real-time, scalable message ingestion
- Dataflow: Unified stream and batch processing with custom logic
- BigQuery: Petabyte-scale data warehousing for analytics
- Cloud Bigtable: Low-latency NoSQL for time-series and operational data
- Handles millions of events/sec, real-time processing, long-term storage
Memory trick: Pub/Sub gets it, Dataflow cooks it, BigQuery archives it, Bigtable shows it.