A data engineering team processes petabytes of streaming data using Amazon Kinesis Data Streams and then stores the raw data in Amazon S3 for long-term retention and future analysis. The data in S3 is accessed very infrequently after the initial processing, but must be available for audit purposes for 10 years. What is the most cost-optimized storage solution for this raw data in S3?
- AAmazon S3 Intelligent-Tiering with a lifecycle policy to move to S3 Glacier Flexible Retrieval after 90 days.
- BAmazon S3 Standard with a lifecycle policy to move to S3 Standard-IA after 30 days.
- CAmazon S3 Standard with a lifecycle policy to transition immediately to S3 Glacier Deep Archive.
- DAmazon S3 Standard with a lifecycle policy to move to S3 Glacier Flexible Retrieval after 30 days and then to S3 Glacier Deep Archive after 90 days.
Show answer & explanationAnswer & explanation
Correct answer: D. Amazon S3 Standard with a lifecycle policy to move to S3 Glacier Flexible Retrieval after 30 days and then to S3 Glacier Deep Archive after 90 days.
The requirement states that data is accessed 'very infrequently after initial processing' and needs 'long-term retention' (10 years). A lifecycle policy that first moves data to S3 Glacier Flexible Retrieval after 30 days (for potential infrequent access with flexible retrieval) and then to S3 Glacier Deep Archive after 90 days (for the lowest cost, long-term archival) provides the most cost-optimized solution while meeting the retention and access pattern needs.
Why the other options are wrong
- A. Intelligent-Tiering is good for unknown access patterns, but here the pattern is known ('very infrequently accessed'). A direct transition to Glacier classes after initial processing is more cost-effective. The specific lifecycle policy here might be less optimal than a multi-stage Glacier transition.
- B. Moving to S3 Standard-IA is good for infrequent access but not the most cost-effective for 'very infrequently accessed' data over 10 years, as Glacier classes are cheaper for deep archives.
- C. Immediately transitioning to S3 Glacier Deep Archive might be too aggressive if there's any chance of access in the first 30-90 days, as Deep Archive has higher retrieval costs and longer retrieval times. A staged approach is often more flexible and cost-effective.
S3 Lifecycle Policies for Cost Optimization
Amazon S3 Lifecycle policies allow you to automatically transition objects to different S3 storage classes or expire them after a defined period, optimizing storage costs based on access patterns.
- Automate movement between S3 Standard, Standard-IA, One Zone-IA, Glacier Instant Retrieval, Glacier Flexible Retrieval, and Glacier Deep Archive.
- Reduce storage costs by moving older, less frequently accessed data to cheaper storage classes.
- Can also be used to automatically delete objects after a certain period.
- Transitions incur a small per-object charge.
Memory trick: S3 lifecycle is like a data conveyor belt, moving items to cheaper storage as they age.