AWS Certified Machine Learning – SpecialtyData EngineeringEasy
A data engineering team is building a data lake on Amazon S3 for an analytics platform. They receive data from various sources in different formats (CSV, JSON, Parquet). To enable efficient querying and machine learning model training without rewriting queries for each format, they need a centralized metadata repository that describes the schema, location, and partition information of all datasets in the data lake. This repository should also allow integration with services like Amazon Athena and AWS Glue. Which AWS service is best suited for this purpose?
- AAmazon RDS
- BAmazon DynamoDB
- CAWS Glue Data Catalog
- DAmazon Redshift
Show answer & explanationAnswer & explanation
Correct answer: C. AWS Glue Data Catalog
The AWS Glue Data Catalog is a fully managed, persistent metadata store that stores schema, location, and partition information for data in Amazon S3 and other data stores. It is central to building a data lake and integrates seamlessly with services like Amazon Athena, AWS Glue ETL, and Amazon Redshift Spectrum, enabling consistent querying across various data formats.
Why the other options are wrong
- A. Amazon RDS is a relational database service, not suitable for managing metadata across diverse data lake formats.
- B. Amazon DynamoDB is a NoSQL database service, not designed as a metadata catalog for data lakes.
- D. Amazon Redshift is a data warehousing service, not a metadata catalog. While it can query S3 data via Redshift Spectrum, it relies on the Glue Data Catalog for metadata.
AWS Glue Data Catalog
A fully managed, persistent metadata store that acts as a central repository for schema, location, and partition information of data assets in a data lake.
- Serverless metadata management
- Integrates with Athena, Glue ETL, Redshift Spectrum
- Supports various data formats (CSV, JSON, Parquet)
- Enables unified view of data lake data
Memory trick: Glue's catalog is the map to your data lake's schema trap.