AWS Certified Machine Learning – SpecialtyData EngineeringEasy

A data engineering team is building a data lake for a new machine learning project. They need to ensure that the data stored in Amazon S3 is discoverable, queryable, and has a centralized metadata catalog for various AWS analytics and ML services. This catalog should allow different teams to understand the schema and easily query the data without needing to know its underlying physical location or format. Which AWS service should be used to create and manage this metadata catalog?

  1. AAWS Glue Data Catalog
  2. BAmazon EMR
  3. CAmazon DynamoDB
  4. DAmazon Redshift Spectrum
Show answer & explanation

Correct answer: A. AWS Glue Data Catalog

The AWS Glue Data Catalog is a centralized metadata repository for all your data assets across various AWS services. It stores schema information, table definitions, and physical locations of data, making it discoverable and queryable by services like Amazon Athena, Amazon Redshift Spectrum, and Amazon SageMaker.

Why the other options are wrong

  • B. Amazon EMR is a managed cluster platform for running big data frameworks like Apache Spark, Hadoop, and Presto. While it can process data in a data lake, it does not provide the centralized metadata catalog service itself.
  • C. Amazon DynamoDB is a NoSQL database service, not a metadata catalog for a data lake. It stores operational data, not schema information about data in S3.
  • D. Amazon Redshift Spectrum allows Redshift to query data directly in S3 using external tables defined in the Glue Data Catalog. It uses the catalog but is not the catalog itself.

AWS Glue Data Catalog

A persistent metadata store that stores schema and table definitions for data in your data lake, making it discoverable and queryable by other AWS services.

  • Centralized metadata repository.
  • Integrates with S3, Athena, Redshift Spectrum, SageMaker, EMR, etc.
  • Stores table definitions, schemas, partition information.
  • Crucial for building a functional data lake.
  • Managed and serverless.

Memory trick: Glue's Catalog holds the map to your data lake treasures.

More Data Engineering questions