Professional Data EngineerEnsuring solution qualityHard
A global ride-sharing company is building a new data processing pipeline to analyze driver and rider location data. Due to the massive scale (trillions of records) and the need for extremely low-latency queries (milliseconds) for real-time decision-making (e.g., dynamic pricing, driver matching), a traditional relational database or standard data warehouse is insufficient. The data is primarily time-series, with new data constantly appended and historical data frequently accessed. You need to select a Google Cloud database service that can handle this scale and performance requirement. Which service is most appropriate?
- ACloud Spanner.
- BCloud SQL.
- CCloud Bigtable.
- DBigQuery.
Show answer & explanationAnswer & explanation
Correct answer: C. Cloud Bigtable.
Cloud Bigtable is a fully managed, NoSQL wide-column database service designed for massive scale (petabytes of data) and extremely low-latency reads and writes (milliseconds). It is particularly well-suited for time-series data, operational analytics, and high-throughput applications where real-time performance on large datasets is critical, making it ideal for the described scenario.
Why the other options are wrong
- A. Cloud Spanner is a globally distributed, strongly consistent relational database, offering high availability and scalability, but its primary use case is transactional workloads requiring strong consistency, not typically for 'trillions of records' of time-series data with millisecond operational analytical queries, where Bigtable excels.
- B. Cloud SQL is a relational database and cannot handle 'trillions of records' with 'millions of events per second' and 'milliseconds' latency for operational queries.
- D. BigQuery is a highly scalable data warehouse for analytical queries, but its latency is typically in seconds, not milliseconds, and it's not optimized for operational, record-level lookups at this scale with millisecond latency.
Cloud Bigtable for Time-Series Data
A fully managed, NoSQL wide-column database optimized for massive-scale time-series, operational, and analytical workloads requiring low-latency reads/writes.
- Massive scale (petabytes).
- Extremely low latency (milliseconds).
- Ideal for time-series and operational analytics.
- High throughput for reads and writes.
Memory trick: For trillions of records and millisecond queries, Bigtable is your giant, super-fast spreadsheet for real-time decisions.