AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy

A research team needs to process petabytes of scientific simulation data stored in Amazon S3. The data is in various formats, including CSV, JSON, and Parquet. They require a serverless, ad-hoc query engine that can directly query data in S3 without needing to load it into a database, supporting standard SQL for data exploration and analysis. Which AWS service is best suited for this requirement?

  1. AAmazon Athena
  2. BAmazon Redshift
  3. CAWS Glue ETL
  4. DAmazon DynamoDB
Show answer & explanation

Correct answer: A. Amazon Athena

Amazon Athena is a serverless interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It's ideal for ad-hoc querying of various file formats without provisioning servers.

Why the other options are wrong

  • B. Amazon Redshift is a fully managed, petabyte-scale data warehouse. It requires data to be loaded into it and is not serverless for ad-hoc querying directly on S3 data.
  • C. AWS Glue ETL is a serverless data integration service for ETL jobs. While it can transform data, it's not primarily an ad-hoc query engine for direct analysis of S3 data using SQL.
  • D. Amazon DynamoDB is a NoSQL key-value and document database. It's not designed for ad-hoc SQL querying of data stored in S3.

Amazon Athena

A serverless interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL.

  • Serverless, pay-per-query pricing.
  • Queries data directly in S3, no data loading.
  • Supports various data formats (CSV, JSON, Parquet, ORC).

Memory trick: Athena queries S3 data, serverless and free.

More Data Ingestion and Transformation questions