AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A data engineering team is designing a data lake for a global e-commerce platform. They need to store petabytes of customer order data, product catalog information, and website clickstream logs. The data will be accessed by various analytics teams and machine learning models, with varying access patterns. Some teams require real-time access to recent order data, while others perform historical trend analysis on older, less frequently accessed data. Cost optimization is a major concern. Which Amazon S3 storage class strategy should they implement to balance performance and cost for this diverse workload?

  1. AStore recent order data in S3 Standard-IA, clickstream logs in S3 Intelligent-Tiering, and product catalog in S3 Standard.
  2. BStore all frequently accessed data in S3 Standard and less frequently accessed data in S3 Standard-IA, with S3 Lifecycle policies to transition to S3 Glacier Flexible Retrieval for historical data.
  3. CStore all data in S3 One Zone-IA to minimize costs, and create replicas in another region for disaster recovery.
  4. DStore all data in S3 Standard and use S3 Lifecycle policies to transition to S3 Glacier Deep Archive after 90 days.
Show answer & explanation

Correct answer: B. Store all frequently accessed data in S3 Standard and less frequently accessed data in S3 Standard-IA, with S3 Lifecycle policies to transition to S3 Glacier Flexible Retrieval for historical data.

This strategy effectively balances cost and performance. S3 Standard is ideal for frequently accessed data like recent orders. S3 Standard-IA is cost-effective for less frequently accessed data that still requires rapid access, suitable for some analytics. S3 Glacier Flexible Retrieval (formerly S3 Glacier) is suitable for archival data that is rarely accessed but needs flexible retrieval options, making it perfect for historical trend analysis. S3 Lifecycle policies automate the transitions, optimizing costs over time.

Why the other options are wrong

  • A. S3 Standard-IA for recent order data might be too slow if real-time access is truly needed. S3 Intelligent-Tiering is good for unknown access patterns, but specifying Standard for product catalog and Standard-IA for recent orders doesn't fully optimize for diverse, known access patterns and historical archives.
  • C. S3 One Zone-IA offers lower cost but lacks multi-AZ redundancy, making it less resilient for critical data. Replicas in another region would significantly increase costs and complexity, negating the cost advantage for the primary storage class choice.
  • D. S3 Glacier Deep Archive has the lowest cost but highest retrieval times (hours to days), which is unsuitable for any data requiring even occasional access for trend analysis.

S3 Storage Class Optimization

Selecting the most appropriate Amazon S3 storage class for data based on access patterns, durability requirements, and cost objectives to optimize storage spend.

  • S3 Standard: Frequent access, high performance.
  • S3 Standard-IA: Infrequent access, rapid retrieval.
  • S3 Intelligent-Tiering: Automatic tiering for changing/unknown access.
  • S3 Glacier Flexible Retrieval: Archival, flexible retrieval options.
  • S3 Glacier Deep Archive: Lowest cost archival, longest retrieval times.
  • S3 One Zone-IA: Infrequent access, single availability zone.

Memory trick: Match your data's 'hotness' to S3's 'coolness' for optimal savings.

More Data Engineering questions