A global marketing analytics company needs to process clickstream data from millions of users across various websites and mobile apps. The data is generated continuously, and they need to perform near real-time aggregations (e.g., hourly unique visitors, page views per campaign) to feed into live dashboards. Additionally, this processed data must be stored for long-term historical analysis and machine learning model training. The solution must be cost-effective and highly available. Which approach should be used to build this data pipeline?
- AUse Pub/Sub for ingestion, Dataflow for streaming processing, and BigQuery for both real-time aggregations and historical storage.
- BUse Cloud SQL for raw data, Cloud Functions for processing, and Cloud Spanner for analytics.
- CUse Cloud Storage for raw data, Dataflow for batch processing, and BigQuery for analytics.
- DUse Cloud Storage for raw data, custom Python scripts on Compute Engine for processing, and Cloud Bigtable for analytics.
Show answer & explanationAnswer & explanation
Correct answer: A. Use Pub/Sub for ingestion, Dataflow for streaming processing, and BigQuery for both real-time aggregations and historical storage.
This scenario requires a real-time streaming pipeline for near real-time aggregations and a scalable data warehouse for long-term storage. Pub/Sub handles high-volume, continuous ingestion. Dataflow in streaming mode performs the near real-time aggregations. BigQuery can efficiently store both the aggregated data for live dashboards and the raw/processed data for historical analysis and ML training, offering scalability and cost-effectiveness.
Why the other options are wrong
- B. Cloud SQL and Cloud Spanner are not optimized for high-volume, continuous clickstream data ingestion and large-scale analytical processing. Cloud Functions are for short-lived tasks.
- C. Batch processing with Dataflow would not meet the 'near real-time' requirement for live dashboards.
- D. Custom scripts on Compute Engine introduce significant operational overhead and are less efficient than managed services like Dataflow. Cloud Bigtable is for operational NoSQL, not primary for analytical dashboards and ML training on aggregated data.
Streaming Analytics Pipeline
An architecture designed to ingest, process, and analyze continuous streams of data in near real-time, often combining a message queue, a stream processing engine, and an analytical data store.
- Handles high-volume, continuous data streams.
- Provides low-latency processing and aggregations.
- Supports both real-time dashboards and historical analysis.
- Commonly used for clickstream, IoT, and financial data.
Memory trick: To analyze every click in real-time and keep it forever, you need a data river that flows fast into a huge, smart lake.