AWS Certified Data Engineer – AssociateData Governance and SecurityHard

A data platform team needs to implement a data retention policy for their Amazon S3 data lake. They have varying retention requirements for different datasets, some requiring 3 years, others 5 years, and some 10 years, before permanent deletion. All data is initially stored in S3 Standard. The goal is to automate transitions to a cost-effective infrequent access tier after 30 days and then to an archival tier after 1 year, before final deletion. How should the data engineer configure S3 lifecycle policies to efficiently manage these diverse retention periods and transitions?

  1. AImplement AWS Lambda functions triggered by S3 Events to move objects to different storage classes and delete them based on their age and metadata.
  2. BUse S3 object tagging to categorize datasets by retention period and create a single lifecycle policy with multiple rules, each using a tag filter.
  3. CCreate multiple S3 buckets, one for each retention period, and apply a single lifecycle policy to each bucket.
  4. DCreate a single lifecycle policy with multiple rules, each using a prefix filter for a dataset and defining specific transitions and expiration times.
Show answer & explanation

Correct answer: B. Use S3 object tagging to categorize datasets by retention period and create a single lifecycle policy with multiple rules, each using a tag filter.

Using S3 object tagging to categorize datasets by their required retention period, combined with a single lifecycle policy containing multiple rules (each filtered by a specific tag), allows for flexible and efficient management of diverse retention policies. Each rule can then specify the 30-day transition to Infrequent Access, 1-year transition to Glacier, and the final expiration based on the tag-defined retention.

Why the other options are wrong

  • A. While Lambda could achieve this, it adds operational complexity, custom code to maintain, and potential costs, when S3 lifecycle policies offer a native, managed, and more cost-effective solution for time-based transitions and expirations.
  • C. Creating multiple buckets for different retention periods can lead to bucket sprawl, increased management overhead, and potential misconfigurations. It's generally less efficient than managing within a single bucket using tags or prefixes.
  • D. Prefix filters can work but are less flexible than tags, especially if a dataset's logical grouping doesn't align with its path prefix. Tags offer more granular and dynamic categorization.

S3 Lifecycle Policy with Tag Filters

S3 lifecycle policies can leverage object tags to apply distinct storage class transitions and expiration rules to different groups of objects within the same bucket, offering highly flexible data retention management.

  • Tags provide flexible metadata for categorization
  • Rules apply only to objects matching specified tags
  • Enables diverse retention policies within a single bucket

Memory trick: Tags filter lifecycle rules for diverse retention needs.

More Data Governance and Security questions