AWS Certified Data Engineer – AssociateData Storage and ManagementMedium
A data engineering team manages a large data lake on Amazon S3. They use AWS Glue Data Catalog to store metadata for querying with Amazon Athena. Occasionally, new partitions are added to S3 by an external process, but these new partitions are not immediately discoverable by Athena queries. The team needs a mechanism to automatically update the Glue Data Catalog with these new partitions without manual intervention. Which solution should they implement?
- ARun `ALTER TABLE ADD PARTITION` manually in Athena after each new partition is added.
- BConfigure an S3 event notification to trigger an AWS Lambda function that calls `msck repair table` or `ALTER TABLE ADD PARTITION`.
- CModify the external process to directly update the Glue Data Catalog using the AWS SDK.
- DSchedule a daily AWS Glue Crawler run to discover new partitions.
Show answer & explanationAnswer & explanation
Correct answer: B. Configure an S3 event notification to trigger an AWS Lambda function that calls `msck repair table` or `ALTER TABLE ADD PARTITION`.
Configuring S3 event notifications to trigger a Lambda function is the most automated and real-time solution. The Lambda function can then execute `msck repair table` or directly use Glue API calls to add partitions, ensuring the Data Catalog is updated promptly after new data arrives without waiting for a scheduled crawler or manual intervention.
Why the other options are wrong
- A. Manual `ALTER TABLE ADD PARTITION` is not automated and requires intervention.
- C. Modifying an external process is often complex, introduces tight coupling, and might not be feasible if the external process is outside direct control.
- D. Scheduling a daily Glue Crawler introduces latency and is not 'immediate' discovery, as new partitions might not be available for queries until the next crawl.
Automated Partition Discovery
The process of automatically updating metadata catalogs (like AWS Glue Data Catalog) with new data partitions created in a data lake, ensuring immediate data discoverability for query engines.
- Crucial for data lakes with frequently updated data.
- Prevents manual metadata management overhead.
- Enables query engines to access the latest data.
Memory trick: Events trigger Glue to get new parts.