AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium
A data engineering team needs to build an ETL pipeline for a marketing analytics platform. Raw campaign data arrives daily as CSV files in an Amazon S3 bucket. The team needs to clean, normalize, and enrich this data by joining it with customer segmentation data from an Amazon RDS PostgreSQL database. The transformed data should be stored in Parquet format in another S3 bucket, partitioned by date, and made available for querying by Amazon Athena. The solution needs to be fully managed, serverless, and support complex data transformations without managing underlying compute infrastructure. Which AWS service is the most appropriate for this transformation workload?
- AAWS Step Functions orchestrating AWS Lambda functions
- BAmazon Redshift Spectrum
- CAmazon EMR with Apache Spark
- DAWS Glue ETL
Show answer & explanationAnswer & explanation
Correct answer: D. AWS Glue ETL
AWS Glue ETL is a fully managed, serverless service specifically designed for batch ETL workloads. It can read data from S3 (CSV), connect to RDS PostgreSQL for joins, perform complex transformations using Spark, and write partitioned Parquet data back to S3, automatically updating the Glue Data Catalog for Athena querying. This perfectly matches the requirements for a serverless, managed batch ETL pipeline.
Why the other options are wrong
- A. While Lambda can perform transformations, orchestrating complex ETL involving large datasets and database joins with Step Functions and Lambda becomes cumbersome and less efficient than Glue ETL.
- B. Redshift Spectrum allows querying data in S3 but is not an ETL service; it doesn't perform the cleaning, normalization, or joining with RDS as required.
- C. EMR with Spark can perform these transformations but requires cluster management, which contradicts the 'serverless' requirement.
AWS Glue ETL
A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development.
- Fully managed, serverless Spark-based ETL jobs.
- Integrates with AWS Glue Data Catalog for metadata management.
- Supports various data sources and sinks, including S3 and relational databases.
Memory trick: Glue's serverless Spark cleans, joins, and preps data for analysis.