AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy
A media company needs to process petabytes of video analytics data stored in Amazon S3. The data scientists require the ability to run ad-hoc SQL queries directly on this data without provisioning or managing any servers. The queries often involve complex joins and aggregations across multiple large datasets. The solution must be cost-effective, paying only for the data scanned. Which AWS service is best suited for this requirement?
- AAmazon DynamoDB
- BAmazon Redshift
- CAmazon Athena
- DAWS Glue Data Catalog
Show answer & explanationAnswer & explanation
Correct answer: C. Amazon Athena
Amazon Athena is an interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It is serverless, so there is no infrastructure to manage, and you pay only for the queries you run.
Why the other options are wrong
- A. Amazon DynamoDB is a NoSQL database, not suitable for complex SQL queries on S3 data.
- B. Amazon Redshift is a fully managed data warehouse, requiring provisioning and management, which goes against the 'no provisioning' requirement.
- D. AWS Glue Data Catalog is a metadata repository for ETL operations, not a query engine itself.
Amazon Athena
An interactive query service that makes it easy to analyze data directly in Amazon S3 using standard SQL. It is serverless and pays per query.
- Serverless SQL query service
- Queries data directly in S3
- Pay-per-query pricing model
Memory trick: Athena queries S3 data like a wise owl.