Professional Data EngineerDesigning data processing systemsHard

A large retail company needs to migrate its on-premises operational analytics database to Google Cloud. This database handles millions of transactions per second, requires extremely low-latency reads and writes (single-digit milliseconds), and stores data with a simple key-value structure but needs to support complex aggregations and time-series analysis. The data grows to petabytes. Which Google Cloud service is the most appropriate for this workload?

  1. ACloud Bigtable.
  2. BCloud Spanner.
  3. CCloud SQL.
  4. DBigQuery.
Show answer & explanation

Correct answer: A. Cloud Bigtable.

Cloud Bigtable is a petabyte-scale, fully managed NoSQL database service specifically designed for high-throughput, low-latency workloads. Its wide-column store model is ideal for key-value data with complex aggregations and time-series analysis due to its efficient row key design and column family structure. It can handle millions of operations per second with single-digit millisecond latency, making it perfect for operational analytics at this scale.

Why the other options are wrong

  • B. Cloud Spanner is a globally distributed relational database, offering strong consistency and high availability, but it's not optimized for the raw throughput and low-latency key-value access patterns of Bigtable, especially for time-series and operational analytics at this scale.
  • C. Cloud SQL is a regional relational database and cannot handle millions of transactions per second at petabyte scale with single-digit millisecond latency.
  • D. BigQuery is a data warehouse optimized for analytical queries on structured data, not for operational low-latency reads/writes of millions of transactions per second.

Bigtable for Operational Analytics

A fully managed, petabyte-scale NoSQL wide-column database optimized for high-throughput, low-latency operational analytics, suited for time-series and complex aggregations.

  • Handles millions of ops/sec with low latency.
  • Ideal for time-series and operational data.
  • Scales to petabytes of data.

Memory trick: Bigtable: Big throughput, Big data, Big analytics.

More Designing data processing systems questions