Professional Data EngineerBuilding and operationalizing data processing systemsEasy
A data analytics team is migrating an on-premises Hadoop cluster to Google Cloud. They need to store petabytes of historical log data and frequently accessed reports, ensuring high durability and availability. The data will be accessed by various analytics tools and machine learning workloads. Which storage service should they choose for this data lake?
- ACloud Filestore
- BCloud SQL
- CCloud Storage
- DCloud Spanner
Show answer & explanationAnswer & explanation
Correct answer: C. Cloud Storage
Cloud Storage is an object storage service suitable for data lakes, offering high durability, availability, and scalability for petabytes of unstructured and semi-structured data, making it ideal for analytics and ML workloads.
Why the other options are wrong
- A. Cloud Filestore is a managed file storage service, primarily for applications requiring a file system interface, not a data lake.
- B. Cloud SQL is a managed relational database for structured data, not suitable for petabyte-scale unstructured log data lakes.
- D. Cloud Spanner is a relational database for transactional workloads, not a data lake storage solution.
Google Cloud Storage for Data Lakes
Google Cloud Storage is a highly scalable, durable, and available object storage service, commonly used as the foundation for data lakes on Google Cloud.
- Object storage for any data type
- High durability (11 nines) and availability
- Cost-effective for large volumes of data
Memory trick: Cloud Storage: The vast lake for all your data.