AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy

A data analytics team is using Amazon S3 to store raw log files from various applications. They need to query this data directly in S3 using standard SQL without provisioning or managing any servers. The queries are ad-hoc and exploratory, and the team wants to minimize costs by paying only for the data scanned. Which AWS service should they use for this requirement?

  1. AAmazon RDS PostgreSQL
  2. BAmazon Redshift
  3. CAmazon Athena
  4. DAWS Glue ETL
Show answer & explanation

Correct answer: C. Amazon Athena

Amazon Athena is a serverless interactive query service that enables analysts to query data directly in Amazon S3 using standard SQL. It charges based on the amount of data scanned per query, making it cost-effective for ad-hoc and exploratory analysis without managing any infrastructure.

Why the other options are wrong

  • A. Amazon RDS PostgreSQL is a relational database for transactional workloads, not suited for querying petabytes of raw log files in S3.
  • B. Amazon Redshift is a fully managed data warehouse, which requires provisioning clusters and is optimized for complex, consistent analytical workloads, not ad-hoc querying directly on S3 for cost-efficiency.
  • D. AWS Glue ETL is for data transformation and loading, not for interactive querying of raw data in S3.

Amazon Athena

An interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL.

  • Serverless, so there is no infrastructure to manage.
  • Pay-per-query pricing based on data scanned.
  • Integrates with AWS Glue Data Catalog for schema definition.

Memory trick: Athena: Your SQL lens for S3 data.

More Data Ingestion and Transformation questions