Professional Data EngineerBuilding and operationalizing data processing systemsMedium

A data analytics team is migrating an on-premises Hadoop cluster to Google Cloud. They need to store petabytes of raw, unstructured, and semi-structured data (logs, images, JSON files) for long-term retention and future analysis. The data will be accessed by various analytics tools, including BigQuery, Dataflow, and custom machine learning models. They require a scalable, durable, and cost-effective storage solution that can handle diverse data formats without imposing a rigid schema. Which Google Cloud service is the most appropriate for this data lake requirement?

  1. ACloud Spanner
  2. BCloud SQL
  3. CCloud Storage
  4. DBigQuery
Show answer & explanation

Correct answer: C. Cloud Storage

Cloud Storage is an object storage service that is highly scalable, durable, cost-effective, and schema-agnostic, making it the ideal choice for building a data lake to store diverse raw data for various downstream analytics and ML workloads.

Why the other options are wrong

  • A. Cloud Spanner is a globally distributed relational database, overkill and unsuitable for raw data lake storage.
  • B. Cloud SQL is a relational database and not suitable for storing petabytes of raw, unstructured, and semi-structured data in a data lake.
  • D. BigQuery is an analytical data warehouse designed for structured data and SQL queries, not for raw, diverse data lake storage.

Cloud Storage for Data Lakes

Utilizing Google Cloud Storage as the foundational layer for a data lake due to its scalability, durability, cost-effectiveness, and ability to store any data format.

  • Object storage, schema-agnostic
  • Unlimited scalability
  • Integration with various GCP analytics services (Dataflow, BigQuery, Dataproc)

Memory trick: Cloud Storage holds all the lake's treasures.

More Building and operationalizing data processing systems questions