Professional Cloud ArchitectDesign and plan a cloud solution architectureHard

A media company hosts a popular online streaming platform on Google Cloud. They need to analyze user behavior data, including clickstreams, viewing history, and ratings, to provide personalized recommendations. This data arrives continuously and at high velocity, requiring immediate processing to update user profiles and recommendation models. The solution must handle petabytes of data, scale elastically, and provide low-latency query capabilities for real-time personalization. Which data architecture best meets these requirements?

  1. ACloud SQL for transactional data, Cloud Storage for backups, Looker for dashboards.
  2. BCloud Storage for raw data, Dataflow for batch processing, BigQuery for analysis.
  3. CCloud Functions for event handling, Firestore for user profiles, Cloud Spanner for large-scale transactions.
  4. DPub/Sub for ingestion, Dataflow for stream processing, Bigtable for real-time lookups, BigQuery for historical analysis.
Show answer & explanation

Correct answer: D. Pub/Sub for ingestion, Dataflow for stream processing, Bigtable for real-time lookups, BigQuery for historical analysis.

This architecture leverages Pub/Sub for real-time ingestion, Dataflow for scalable stream processing and transformation of continuous data, Bigtable for low-latency lookups of constantly updating user profiles, and BigQuery for comprehensive historical analysis and model training. This combination is ideal for petabyte-scale, high-velocity, real-time personalization.

Why the other options are wrong

  • A. Cloud SQL is for transactional relational data, not petabyte-scale streaming analytics. Cloud Storage for backups and Looker for dashboards are relevant but don't address the core processing and storage for real-time personalization at this scale.
  • B. This is a batch processing architecture, unsuitable for 'continuously and at high velocity' data requiring 'immediate processing' and 'real-time personalization'.
  • C. Cloud Functions are too limited for continuous petabyte-scale stream processing. Firestore is good for user profiles but might struggle with petabyte scale and high-velocity updates for real-time lookups compared to Bigtable. Cloud Spanner is for transactional relational data, not analytics and real-time lookups of profile data.

Real-time Personalization Architecture

An advanced data architecture designed to ingest, process, and analyze high-velocity, high-volume data streams to provide immediate, personalized user experiences.

  • Uses message queues (Pub/Sub) for ingestion.
  • Employs stream processing (Dataflow) for transformations and real-time updates.
  • Leverages NoSQL databases (Bigtable) for low-latency profile lookups.
  • Integrates with data warehouses (BigQuery) for historical analysis and model training.

Memory trick: Ingest, Stream, Store, Analyze: The four pillars of real-time personalization.

More Design and plan a cloud solution architecture questions