Professional Data EngineerDesigning data processing systemsEasy

A data engineering team is building a new data lake on Google Cloud to store vast amounts of raw, unstructured, and semi-structured data from various sources. The data lake must support flexible schema evolution, integrate with various analytics tools, and provide cost-effective long-term storage. Which Google Cloud service is the most appropriate foundational component for this data lake?

  1. ACloud SQL.
  2. BCloud Storage.
  3. CCloud Bigtable.
  4. DBigQuery.
Show answer & explanation

Correct answer: B. Cloud Storage.

Cloud Storage is the foundational service for building a data lake on Google Cloud. It provides highly durable, available, and cost-effective object storage for any data type, supporting flexible schema evolution (as it stores raw files) and easy integration with other analytics services. It's ideal for storing raw, unstructured, and semi-structured data.

Why the other options are wrong

  • A. Cloud SQL is a relational database, not suitable for storing raw, unstructured data in a data lake.
  • C. Cloud Bigtable is a NoSQL wide-column database for high-throughput, low-latency operational workloads, not a general-purpose data lake storage.
  • D. BigQuery is a data warehouse for structured data analytics, not a raw data lake storage for unstructured data.

Cloud Storage Data Lake

Using Cloud Storage as the primary, cost-effective, and flexible storage layer for raw, unstructured, and semi-structured data in a data lake architecture.

  • Stores objects of any size and type.
  • Offers high durability, availability, and scalability.
  • Provides various storage classes for cost optimization.

Memory trick: Cloud Storage: The Lake's Foundation, holding all the Data's nation.

More Designing data processing systems questions