AWS Certified Data Engineer – AssociateData Storage and ManagementMedium

A data engineering team is building a data lake and needs a centralized metadata repository for all their datasets stored in Amazon S3, Amazon RDS, and Amazon Redshift. They require a service that can automatically discover schemas, track data lineage, and integrate with various AWS analytics services. Which AWS service should they use?

  1. AAWS Secrets Manager
  2. BAmazon CloudWatch
  3. CAWS Glue Data Catalog
  4. DAmazon ES (Elasticsearch Service)
Show answer & explanation

Correct answer: C. AWS Glue Data Catalog

AWS Glue Data Catalog is a fully managed metadata repository that allows you to store, annotate, and share metadata across various AWS analytics and machine learning services. It can automatically discover schemas and track data lineage.

Why the other options are wrong

  • A. AWS Secrets Manager securely stores and manages secrets, not metadata for data lakes.
  • B. Amazon CloudWatch is a monitoring and observability service, not a metadata catalog.
  • D. Amazon ES is a search and analytics engine, not a data catalog.

AWS Glue Data Catalog

The AWS Glue Data Catalog is a persistent metadata store for all your data assets, regardless of where they are located. It allows you to store, annotate, and share metadata, enabling various AWS services to discover and query your data.

  • Centralized metadata repository
  • Integrates with S3, RDS, Redshift, etc.
  • Supports schema discovery (Crawlers)
  • Used by Athena, Redshift Spectrum, EMR, Glue ETL

Memory trick: Glue Guides Global Data.

More Data Storage and Management questions