A global IoT company collects sensor data from millions of devices, generating terabytes of time-series data daily. This data needs to be retained for several years for compliance and long-term analysis, but only the most recent data (last 30 days) requires high-performance, low-latency queries for real-time monitoring dashboards. Older data can be accessed with slightly higher latency but must remain cost-effective. How should this data be stored and queried efficiently on Google Cloud?
- AStore recent data in Cloud Bigtable for low-latency access and older data in BigQuery for cost-effective long-term storage.
- BStore recent data in Cloud SQL for fast access and older data in Cloud Storage with custom indexing.
- CStore all data in BigQuery, using partitioning and clustering for performance.
- DStore all data in Cloud Firestore for flexible querying and retention.
Show answer & explanationAnswer & explanation
Correct answer: A. Store recent data in Cloud Bigtable for low-latency access and older data in BigQuery for cost-effective long-term storage.
This scenario requires a hybrid approach. Cloud Bigtable is optimized for high-throughput, low-latency access to time-series data, making it ideal for the most recent 30 days. BigQuery provides a cost-effective solution for petabyte-scale long-term storage and analytical queries on older data, which can tolerate slightly higher latency. This combination optimizes both performance and cost.
Why the other options are wrong
- B. Cloud SQL is not designed for terabytes of time-series data daily and petabyte-scale retention. Cloud Storage with custom indexing would require significant engineering effort and would likely not meet query performance expectations.
- C. While BigQuery can handle petabytes, using it for the 'most recent 30 days' with strict low-latency requirements for individual lookups might not be as performant as Cloud Bigtable, and continuous ingestion of millions of individual writes can be less efficient than Bigtable for operational use cases.
- D. Cloud Firestore is a document database, not ideal for petabyte-scale time-series data with high ingestion rates and complex analytical queries over long periods.
Hybrid Time-Series Storage
A data architecture that combines different storage solutions to optimize for both low-latency access to recent data and cost-effective, long-term storage of historical data, often involving a fast operational database and an analytical data warehouse.
- Balances performance (low-latency) with cost-effectiveness.
- Typically uses a NoSQL database for recent, 'hot' data.
- Leverages a data warehouse or object storage for 'cold' historical data.
- Common pattern for IoT, monitoring, and financial time-series data.
Memory trick: Time-series data is like a river: the fresh water needs quick access, but the old water can sit in a big, cheap reservoir.