AWS Certified Data Engineer – AssociateData Storage and ManagementEasy
A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time streaming data from IoT devices, batch data from on-premises systems, and semi-structured logs from web applications. The team needs a centralized metadata repository that can automatically discover schemas, track data lineage, and be queried by various analytics services like Amazon Athena and Amazon Redshift Spectrum. Which AWS service should the data engineer use to meet these requirements?
- AAmazon DynamoDB
- BAmazon S3
- CAWS Glue Data Catalog
- DAWS Lake Formation
Show answer & explanationAnswer & explanation
Correct answer: C. AWS Glue Data Catalog
AWS Glue Data Catalog is a fully managed metadata repository that allows for schema discovery, data lineage tracking, and integration with various AWS analytics services. It serves as a central catalog for all data assets in a data lake.
Why the other options are wrong
- A. Amazon DynamoDB is a NoSQL database service, not a metadata catalog for a data lake.
- B. Amazon S3 is an object storage service, not a metadata catalog. While it stores the data for the data lake, it doesn't provide schema discovery or lineage tracking.
- D. AWS Lake Formation builds on the AWS Glue Data Catalog but focuses on security, governance, and permissions management, rather than solely acting as the metadata repository itself.
AWS Glue Data Catalog
A persistent metadata store for all your data assets in a data lake, enabling schema discovery, data lineage, and integration with analytics services.
- Centralized metadata repository
- Automatic schema discovery (crawlers)
- Integrates with Athena, Redshift Spectrum, EMR
- Supports various data formats and sources
Memory trick: Glue's catalog is the library for your data lake, organizing everything.