AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy

A financial institution needs to perform complex, ad-hoc queries on petabytes of historical transaction data stored in Amazon S3. The data is stored in various formats, including CSV, JSON, and Parquet. Analysts require a serverless solution that can query data directly in S3 without loading it into a database, and they need to pay only for the data scanned. Which AWS service is best suited for this requirement?

  1. AAmazon DynamoDB
  2. BAmazon Athena
  3. CAWS Glue Data Catalog with Amazon RDS
  4. DAmazon Redshift
Show answer & explanation

Correct answer: B. Amazon Athena

Amazon Athena is a serverless interactive query service that makes it easy to analyze data in Amazon S3 using standard SQL. It supports various data formats, and you only pay for the queries you run.

Why the other options are wrong

  • A. Amazon DynamoDB is a NoSQL database, not suitable for complex ad-hoc SQL queries on petabytes of historical data in S3.
  • C. AWS Glue Data Catalog helps manage metadata, but Amazon RDS is a relational database service that requires data loading and management, not direct querying of S3 data in a serverless manner.
  • D. Amazon Redshift is a data warehouse that requires loading data and managing clusters, which contradicts the 'query directly in S3' and 'serverless' requirements.

Amazon Athena

An interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It's serverless, so there's no infrastructure to manage.

  • Serverless: no servers to manage.
  • Pay per query: charged for data scanned.
  • Supports various data formats (CSV, JSON, ORC, Parquet, Avro).

Memory trick: Athena asks S3, no servers for me!

More Data Ingestion and Transformation questions