AWS Certified Data Engineer – AssociateData Storage and ManagementEasy
A data engineering team is building a new data lake on AWS. They need to store petabytes of raw, unstructured data from various sources, including application logs, social media feeds, and sensor data. The data needs to be highly available, durable, and cost-effective for long-term storage, with occasional access for analytics. Which AWS service is the most appropriate for this primary storage layer?
- AAmazon EBS
- BAmazon S3
- CAmazon RDS
- DAmazon EFS
Show answer & explanationAnswer & explanation
Correct answer: B. Amazon S3
Amazon S3 is the foundational storage service for data lakes, offering virtually unlimited scalability, high durability, and cost-effectiveness for storing unstructured data. It supports various access patterns, from frequent to archival, making it ideal for raw data storage.
Why the other options are wrong
- A. Amazon EBS provides block storage for EC2 instances, not suitable for a data lake's petabyte-scale, unstructured data.
- C. Amazon RDS is a relational database service, not designed for storing petabytes of unstructured raw data.
- D. Amazon EFS provides scalable file storage for EC2, containers, and serverless, but S3 is more cost-effective and scalable for petabyte-scale raw data lakes.
Amazon S3 for Data Lakes
Amazon S3 (Simple Storage Service) is the primary storage layer for data lakes on AWS, providing highly scalable, durable, and cost-effective object storage.
- Supports virtually unlimited data storage.
- Offers 11 nines of durability.
- Various storage classes for different access patterns and costs.
- Integrates with numerous AWS analytics and machine learning services.
Memory trick: S3, the lake's deep sea, holds data free.