AWS Certified Data Engineer – AssociateData Storage and ManagementMedium

A data engineer is designing a data lake solution on AWS. The data lake will ingest real-time streaming data from various sources and store it in Amazon S3. To ensure data quality and schema enforcement for downstream analytics, the engineer needs a mechanism to validate incoming data against predefined schemas and manage metadata. Which AWS service combination should be used?

  1. AAmazon DynamoDB for metadata management, AWS Lambda for schema validation, and Amazon Kinesis for streaming.
  2. BAmazon EFS for storage, AWS Systems Manager Parameter Store for schema, and Amazon MSK for streaming.
  3. CAmazon Redshift for storage, AWS Lake Formation for permissions, and AWS Data Pipeline for schema enforcement.
  4. DAmazon S3 for storage, AWS Glue Data Catalog for metadata management, and AWS Glue crawlers for schema discovery.
Show answer & explanation

Correct answer: D. Amazon S3 for storage, AWS Glue Data Catalog for metadata management, and AWS Glue crawlers for schema discovery.

AWS Glue Data Catalog acts as a central metadata repository, and AWS Glue Crawlers automatically infer schemas from data in S3, storing them in the catalog. This enables schema validation and governance for data lake analytics.

Why the other options are wrong

  • A. DynamoDB is not a data catalog, and Lambda would require custom code for schema validation, which is less integrated than Glue. Kinesis is for streaming, but the question focuses on schema/metadata management post-ingestion to S3.
  • B. EFS is a file system, not ideal for data lake scale. Parameter Store is for configuration, not a data catalog. MSK is for streaming, but again, the core need is schema/metadata management for S3 data.
  • C. Redshift is a data warehouse, not the primary storage for raw data lake (S3). Data Pipeline is for orchestration, not schema enforcement or metadata management in this context.

AWS Glue Data Catalog

AWS Glue Data Catalog is a persistent metadata store for your data assets, serving as a central repository for table definitions, schema versions, and other metadata for data lakes and data warehousing.

  • Central metadata repository for data lakes
  • Stores table definitions and schemas
  • Integrates with Athena, Redshift Spectrum, EMR
  • Populated by Glue Crawlers or manually

Memory trick: Glue Catalog crawls and keeps data schema true.

More Data Storage and Management questions