AWS Certified Data Engineer – AssociateData Storage and ManagementEasy
A data engineering team is designing a new data lake on AWS. They need to store petabytes of unstructured data, such as images and videos, that will be accessed infrequently after an initial upload period. The solution must be highly durable, cost-effective, and provide options for eventual archival. Which AWS storage service is most appropriate for this requirement?
- AAmazon S3
- BAmazon RDS for PostgreSQL
- CAmazon Redshift
- DAmazon EFS
Show answer & explanationAnswer & explanation
Correct answer: A. Amazon S3
Amazon S3 is the ideal choice for storing petabytes of unstructured data due to its high durability, scalability, and cost-effectiveness, with lifecycle policies for archival. It natively supports various data types including images and videos.
Why the other options are wrong
- B. Amazon RDS is a relational database service, not suitable for petabytes of unstructured data.
- C. Amazon Redshift is a data warehousing service, optimized for structured data analytics, not unstructured object storage.
- D. Amazon EFS is a file storage service for EC2 instances, not designed for petabytes of cost-effective unstructured object storage in a data lake.
Amazon S3 for Data Lakes
Amazon S3 is a highly scalable, durable, and cost-effective object storage service widely used as the foundation for data lakes on AWS.
- Stores unstructured data (images, videos, logs, etc.)
- Offers various storage classes for cost optimization
- Provides high durability and availability
- Integrates with other AWS analytics services
Memory trick: Scalable Storage Solutions are Key for Data Lakes.