AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy
A research team needs to process petabytes of scientific simulation data stored in Amazon S3. The data is in various formats, including CSV, JSON, and Parquet. They require a serverless, ad-hoc query engine that can directly query data in S3 without needing to load it into a database, supporting standard SQL for data exploration and analysis. Which AWS service is best suited for this requirement?
- AAmazon Athena
- BAmazon Redshift
- CAWS Glue ETL
- DAmazon DynamoDB
Show answer & explanationAnswer & explanation
Correct answer: A. Amazon Athena
Amazon Athena is a serverless interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It's ideal for ad-hoc querying of various file formats without provisioning servers.
Why the other options are wrong
- B. Amazon Redshift is a fully managed, petabyte-scale data warehouse. It requires data to be loaded into it and is not serverless for ad-hoc querying directly on S3 data.
- C. AWS Glue ETL is a serverless data integration service for ETL jobs. While it can transform data, it's not primarily an ad-hoc query engine for direct analysis of S3 data using SQL.
- D. Amazon DynamoDB is a NoSQL key-value and document database. It's not designed for ad-hoc SQL querying of data stored in S3.
Amazon Athena
A serverless interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL.
- Serverless, pay-per-query pricing.
- Queries data directly in S3, no data loading.
- Supports various data formats (CSV, JSON, Parquet, ORC).
Memory trick: Athena queries S3 data, serverless and free.