Professional Cloud ArchitectDesign and plan a cloud solution architectureMedium
A global manufacturing company needs to collect and process large volumes of real-time sensor data from its factories worldwide. This data needs to be ingested continuously, transformed, and then stored in a data warehouse for analytics. The solution must be highly scalable, fault-tolerant, and support both streaming and batch processing paradigms. Which Google Cloud services should be used to build this data pipeline?
- ACloud SQL, Cloud Functions, Looker
- BCloud IoT Core, Cloud Spanner, Cloud Data Studio
- CCloud Storage, Cloud Dataproc, BigQuery
- DCloud Pub/Sub, Cloud Dataflow, BigQuery
Show answer & explanationAnswer & explanation
Correct answer: D. Cloud Pub/Sub, Cloud Dataflow, BigQuery
Cloud Pub/Sub is ideal for ingesting real-time streaming data. Cloud Dataflow provides a fully managed service for executing Apache Beam pipelines, supporting both streaming and batch processing for data transformation. BigQuery is the scalable data warehouse for analytics.
Why the other options are wrong
- A. Cloud SQL is a transactional database, not for large-scale streaming data ingestion. Cloud Functions are for event-driven computing, not a full-fledged data processing engine, and Looker is for BI, not data storage/transformation.
- B. Cloud IoT Core handles device connections, but Cloud Spanner is a transactional database, not a data warehouse. Cloud Data Studio (now Looker Studio) is for visualization, not data processing or storage.
- C. Cloud Storage is for raw storage, and Cloud Dataproc is for Spark/Hadoop, which could work but Dataflow is often preferred for managed stream processing. This combination is less optimized for real-time ingestion and continuous transformation than Pub/Sub and Dataflow.
Streaming Data Pipeline
A series of interconnected services designed to ingest, process, and analyze data in real-time as it is generated, typically involving messaging, processing, and storage components.
- Ingestion: Pub/Sub for high-throughput messaging.
- Processing: Dataflow for unified streaming/batch transformations.
- Storage: BigQuery for scalable analytical data warehousing.
Memory trick: Pub/Sub gets the stream, Dataflow processes the flow, BigQuery holds the big data.