AWS Certified Machine Learning – SpecialtyData EngineeringEasy
A data engineering team is building a serverless data lake. They ingest data from various sources into Amazon S3. To enable efficient querying by services like Amazon Athena and Amazon Redshift Spectrum, they need a centralized metadata repository that automatically discovers schema and partitions. They also require the ability to define and enforce fine-grained access control on tables and columns without managing separate databases. Which AWS service is purpose-built to meet these requirements?
- AAmazon DynamoDB
- BAmazon EMR Metastore
- CAWS Glue Data Catalog
- DAmazon RDS
Show answer & explanationAnswer & explanation
Correct answer: C. AWS Glue Data Catalog
The AWS Glue Data Catalog is a fully managed, centralized metadata repository for all your data assets, including those in S3. It automatically discovers schemas and partitions using crawlers and allows for fine-grained access control through AWS Lake Formation, making it ideal for serverless data lakes and integration with Athena/Redshift Spectrum.
Why the other options are wrong
- A. Amazon DynamoDB is a NoSQL database, not a metadata catalog for S3 data lakes.
- B. Amazon EMR Metastore can be used with EMR, but it is not a fully managed, serverless, centralized catalog for a broader data lake architecture that integrates with services like Athena without an active EMR cluster.
- D. Amazon RDS is a relational database service, not suitable as a metadata catalog for S3 data lakes.
AWS Glue Data Catalog
A persistent, metadata store for data assets in your data lake, providing schema, table, and partition information for various AWS analytics services.
- Fully managed and serverless.
- Automatically discovers schema and partitions using crawlers.
- Integrates with Amazon Athena, Redshift Spectrum, EMR, and AWS Lake Formation.
- Enables fine-grained access control via Lake Formation.
Memory trick: Glue Catalog knows all, for Athena and Redshift, it's the central hall!