AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium
A data engineering team needs to build an ETL pipeline for a marketing analytics platform. The pipeline must process large volumes of historical customer interaction data (several terabytes) stored in Amazon S3, apply complex business logic (e.g., deduplication, aggregation, sentiment analysis using external libraries), and then load the refined data into an Amazon Redshift data warehouse. The solution needs to be scalable, cost-effective, and provide a serverless approach for ETL job execution. Which AWS service is most appropriate for the transformation phase?
- AAmazon Kinesis Data Analytics
- BAmazon EMR with Apache Spark
- CAWS Glue ETL
- DAWS Data Pipeline
Show answer & explanationAnswer & explanation
Correct answer: C. AWS Glue ETL
AWS Glue ETL is a serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. It is ideal for batch processing of large datasets in S3, applying complex transformations, and loading into Redshift without managing servers.
Why the other options are wrong
- A. Amazon Kinesis Data Analytics is for real-time streaming analytics, not batch ETL of historical S3 data.
- B. Amazon EMR with Apache Spark can perform the required transformations but is not serverless; it requires managing EMR clusters, which goes against the 'serverless approach' requirement.
- D. AWS Data Pipeline is an older orchestration service and does not provide an integrated, serverless ETL engine like Glue. It often requires EC2 instances for actual data processing.
AWS Glue ETL
A serverless data integration service that simplifies the process of discovering, preparing, and combining data for analytics, machine learning, and application development.
- Serverless ETL service
- Supports large batch data transformations
- Integrates with S3, Redshift, and other data stores
Memory trick: Glue transforms batches without a server crew.