Professional Data EngineerDesigning data processing systemsMedium
A data analytics team needs to build a robust data pipeline to ingest real-time sensor data from thousands of IoT devices. The data needs to be processed with minimal latency, enriched with contextual information, and then stored in a data warehouse for further analysis. The system must be highly scalable to handle fluctuating data volumes and resilient to device failures or network outages. Which set of Google Cloud services should be used to design this pipeline?
- ACloud IoT Core, Compute Engine, Cloud SQL.
- BCloud Pub/Sub, Dataflow, BigQuery.
- CCloud Functions, Cloud Spanner, Looker Studio.
- DCloud Storage, Dataproc, BigQuery.
Show answer & explanationAnswer & explanation
Correct answer: B. Cloud Pub/Sub, Dataflow, BigQuery.
This scenario describes a classic real-time streaming pipeline. Cloud Pub/Sub is ideal for ingesting high-volume, real-time data from IoT devices. Dataflow (Apache Beam) provides a powerful, managed service for stream processing, allowing for low-latency enrichment and transformations. BigQuery is the suitable data warehouse for storing and analyzing the processed data at scale.
Why the other options are wrong
- A. Cloud IoT Core is for device management, but Compute Engine requires manual scaling and management, and Cloud SQL is not suitable for petabyte-scale data warehousing for analysis.
- C. Cloud Functions are for event-driven, short-lived tasks, not continuous stream processing. Cloud Spanner is a transactional database, not a data warehouse. Looker Studio is for visualization, not data storage or processing.
- D. Cloud Storage and Dataproc are primarily for batch processing and data lake scenarios, not real-time stream ingestion and processing.
Real-time Streaming Pipeline
A data pipeline designed to ingest, process, and analyze data in real-time, often for immediate insights or operational decisions.
- Uses Pub/Sub for scalable ingestion.
- Leverages Dataflow for stream processing.
- Stores processed data in BigQuery for analytics.
Memory trick: Pub/Sub Pushes, Dataflow Digests, BigQuery's Best for Data Storage.