Google Cloud Digital LeaderGeneral knowledge of Google CloudMedium
A data analytics team needs to process petabytes of streaming data from IoT devices in real-time and then run complex SQL queries on the aggregated data for business intelligence. They prefer a serverless solution to minimize operational overhead. Which two Google Cloud products are best suited for this end-to-end scenario?
- ACloud Bigtable and Compute Engine
- BCloud Storage and Dataflow
- CPub/Sub and BigQuery
- DCloud SQL and Data Studio
Show answer & explanationAnswer & explanation
Correct answer: C. Pub/Sub and BigQuery
Pub/Sub is a serverless messaging service ideal for ingesting high volumes of streaming data from IoT devices. BigQuery is a serverless, highly scalable data warehouse optimized for running complex SQL queries on petabytes of data, perfect for subsequent business intelligence. Together, they form a common serverless streaming analytics pipeline.
Why the other options are wrong
- A. Cloud Bigtable is a NoSQL database for large analytical/operational workloads, not primarily for SQL queries. Compute Engine requires managing VMs.
- B. Cloud Storage is for object storage, not real-time streaming ingestion. Dataflow is for processing, but needs an ingestion service.
- D. Cloud SQL is a relational database (not for petabyte-scale streaming analytics). Data Studio is a visualization tool, not a data processing or storage service.
Google Cloud Streaming Analytics Pipeline
Commonly uses Pub/Sub for real-time ingestion and BigQuery for serverless, petabyte-scale SQL analytics on streaming data.
- Pub/Sub acts as a scalable message queue for ingress.
- BigQuery provides a serverless SQL data warehouse for analysis.
- Often combined with Dataflow for complex transformations between ingestion and storage.
Memory trick: Stream with Pub/Sub, Query with BigQuery, Visualize with Looker.