Professional Data EngineerDesigning data processing systemsMedium
A data engineering team is designing a data processing system for a global e-commerce platform. They need to store and analyze customer order history, which can grow to petabytes in size, and perform complex analytical queries on this data for business intelligence and reporting. The solution must provide high availability and strong consistency for query results, while also being cost-effective for long-term storage and ad-hoc analysis. Which Google Cloud product is the most appropriate choice for this scenario?
- ABigQuery
- BCloud SQL
- CCloud Spanner
- DCloud Bigtable
Show answer & explanationAnswer & explanation
Correct answer: A. BigQuery
BigQuery is a fully managed, serverless enterprise data warehouse that can handle petabytes of data. It is optimized for complex analytical queries, provides high availability, and offers cost-effective storage with a pay-as-you-go model, making it ideal for the described scenario.
Why the other options are wrong
- B. Cloud SQL is a managed relational database for OLTP workloads, not designed for petabyte-scale analytical warehousing.
- C. Cloud Spanner is a globally distributed relational database, optimized for transactional workloads, not primarily for petabyte-scale analytical warehousing.
- D. Cloud Bigtable is a NoSQL wide-column database, designed for high-throughput, low-latency operational applications, not typically for complex SQL analytics on petabytes of historical data.
BigQuery for Petabyte-scale Data Warehousing
BigQuery is a serverless, highly scalable, and cost-effective enterprise data warehouse designed for petabyte-scale data analytics and business intelligence.
- Handles petabytes of data.
- Optimized for complex analytical SQL queries.
- Fully managed and serverless, eliminating infrastructure management.
Memory trick: BigQuery: Big Data, Big Queries, Big Savings.