AWS Certified Data Engineer – AssociateData Governance and SecurityHard
A media company stores petabytes of video content in an S3 data lake. This content is accessed by various internal teams (e.g., editorial, marketing, legal) with different access levels. The company wants to implement a solution that allows each team to only view the content relevant to their roles, including specific video clips or segments, without duplicating data. Which AWS service and feature combination should the data engineer use?
- AAmazon Macie for data classification and S3 Access Points
- BAWS Lake Formation with row-level and column-level security
- CAWS Identity and Access Management (IAM) policies with resource-based access
- DS3 Bucket Policies combined with S3 Object Tagging
Show answer & explanationAnswer & explanation
Correct answer: B. AWS Lake Formation with row-level and column-level security
AWS Lake Formation with its fine-grained access controls, including row-level and column-level security (which can be extended to cell-level filtering for specific segments), is designed for precisely this scenario in a data lake. It allows different users or groups to see different subsets of data within the same S3 objects based on their permissions.
Why the other options are wrong
- A. Amazon Macie is for data discovery and classification, not for enforcing access control. S3 Access Points simplify access to S3 buckets but don't provide row/column-level filtering.
- C. IAM policies, while powerful for resource access, don't inherently provide row or column-level filtering within data stored in S3 without additional services or complex logic.
- D. S3 Bucket Policies and Object Tagging can control access at the object level, but not typically at the sub-object (row/column/segment) level without complex and inefficient data duplication or processing.
Lake Formation Fine-Grained Access
AWS Lake Formation provides granular access control for data lakes, enabling row-level, column-level, and cell-level security to filter data based on user identity or attributes without duplicating data.
- Centralized access management for S3 data lakes.
- Filters data at query time.
- Supports various data access patterns.
- Simplifies compliance with data privacy regulations.
Memory trick: Lake Forms Fine-Grained Views