AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A media company needs to process petabytes of video analytics data stored in Amazon S3. The data is in a variety of complex nested JSON and Parquet formats, and analysts require interactive query capabilities without loading data into a traditional database. The solution must support schema evolution and be cost-effective for ad-hoc queries on large datasets. Which AWS service is most appropriate for transforming and querying this data?

  1. AAmazon EMR with Apache Hive
  2. BAWS Glue Data Catalog with Amazon Athena
  3. CAmazon DynamoDB
  4. DAmazon Redshift
Show answer & explanation

Correct answer: B. AWS Glue Data Catalog with Amazon Athena

AWS Glue Data Catalog stores metadata for data in S3, including schema definitions for various formats like JSON and Parquet, supporting schema evolution. Amazon Athena, a serverless query service, can then directly query this data in S3 using standard SQL, making it cost-effective for ad-hoc queries without managing infrastructure.

Why the other options are wrong

  • A. EMR with Hive can query S3, but it requires managing clusters and is less cost-effective and serverless than Athena for ad-hoc queries.
  • C. DynamoDB is a NoSQL database for transactional workloads, not suitable for petabyte-scale analytical queries on S3 data lakes.
  • D. Redshift is a data warehouse requiring data loading and management, not ideal for interactive queries directly on S3 data lakes without loading.

Amazon Athena

A serverless interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL.

  • Pay per query, based on data scanned.
  • Works directly with data in S3, no data loading required.
  • Integrates with AWS Glue Data Catalog for schema management.

Memory trick: Athena queries S3 data, Glue holds the map.

More Data Ingestion and Transformation questions