AWS Certified Data Engineer – AssociateData Ingestion and TransformationEasy
A data analytics team needs to process semi-structured log data (JSON format) generated by various microservices and stored in an Amazon S3 bucket. They require a serverless solution to transform this data into a columnar format (Parquet) and partition it by date for optimized querying. The processing should be scheduled to run daily. Which AWS service is the MOST suitable for this transformation task?
- AAmazon Kinesis Data Firehose
- BAmazon EMR
- CAmazon Athena
- DAWS Glue ETL
Show answer & explanationAnswer & explanation
Correct answer: D. AWS Glue ETL
AWS Glue ETL is a serverless data integration service that excels at reading data from S3, transforming it (e.g., JSON to Parquet, partitioning), and writing it back to S3. It can be scheduled to run daily, making it ideal for this batch transformation scenario.
Why the other options are wrong
- A. Amazon Kinesis Data Firehose is for real-time streaming data ingestion and delivery, not for scheduled batch transformation of existing S3 data.
- B. Amazon EMR is a managed cluster platform for big data processing (e.g., Spark, Hadoop). While it can perform the transformations, it's not serverless; users need to manage cluster lifecycle or use EMR Serverless, but Glue ETL is a more direct and fully managed serverless option for this specific task.
- C. Amazon Athena is a query service for S3 data. While it can query JSON, it's not designed for scheduled transformations to Parquet and partitioning for optimized storage.
AWS Glue ETL
A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development.
- Serverless and scalable ETL jobs.
- Supports various data sources and formats.
- Integrates with Glue Data Catalog for metadata management.
Memory trick: Glue ETL cleans and organizes S3 data, serverless and grand.