Professional Cloud ArchitectManage and provision solution infrastructureMedium

A manufacturing company is collecting sensor data from thousands of IoT devices globally. This data needs to be ingested in real-time, processed for anomalies, and stored for long-term analysis. The solution must be highly scalable, handle unpredictable spikes in data volume, and integrate seamlessly with other Google Cloud analytics services. Which architecture should they implement for the ingestion and initial processing of this streaming data?

  1. ACloud Storage for ingestion, followed by batch processing with Dataproc.
  2. BCloud SQL for ingestion, with Cloud Functions for anomaly detection.
  3. CCloud VPN to on-premises Kafka, then manual data transfer to BigQuery.
  4. DCloud Pub/Sub for ingestion, followed by real-time processing with Cloud Dataflow.
Show answer & explanation

Correct answer: D. Cloud Pub/Sub for ingestion, followed by real-time processing with Cloud Dataflow.

Cloud Pub/Sub provides a highly scalable, global messaging service for real-time data ingestion, handling unpredictable spikes. Cloud Dataflow is ideal for real-time stream processing, including anomaly detection, and integrates well with Pub/Sub and other analytics services.

Why the other options are wrong

  • A. Cloud Storage and Dataproc are primarily for batch processing, not real-time ingestion and processing of streaming data.
  • B. Cloud SQL is a relational database, not optimized for high-volume, real-time streaming ingestion, and Cloud Functions are better for short-lived, event-driven tasks, not continuous stream processing.
  • C. This involves on-premises components and manual steps, which would not be scalable, real-time, or leverage Google Cloud's managed services effectively for this use case.

IoT Data Pipeline (GCP)

A common architecture on Google Cloud for ingesting, processing, and analyzing real-time data from IoT devices.

  • Often starts with Pub/Sub for ingestion.
  • Uses Dataflow for real-time stream processing.
  • Integrates with BigQuery for analytics and Cloud Storage for raw data.

Memory trick: Pub/Sub gets data, Dataflow processes the flow.

More Manage and provision solution infrastructure questions