AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium
A retail company collects point-of-sale (POS) data from thousands of stores. Each store sends small, infrequent batches of sales data (CSV files, ~10-20MB each) to a central location. This data needs to be ingested into an Amazon S3 staging bucket, then transformed and loaded into an Amazon Redshift data warehouse nightly. The transformation involves simple data cleansing, aggregation, and lookup against a product catalog in S3. The solution should be cost-effective for batch processing and easy to manage. Which AWS service is best suited for the transformation and loading (ETL) phase?
- AAWS Glue ETL
- BAmazon Kinesis Data Firehose
- CAmazon EMR
- DAWS Glue DataBrew
Show answer & explanationAnswer & explanation
Correct answer: A. AWS Glue ETL
AWS Glue ETL is a fully managed, serverless Spark-based service ideal for batch ETL workloads like this. It can read CSV from S3, perform cleansing, aggregations, and lookups (even against other S3 data), and then load into Redshift. It is cost-effective for periodic batch processing and requires no server management.
Why the other options are wrong
- B. Kinesis Data Firehose is for streaming ingestion, not for batch ETL transformations.
- C. Amazon EMR provides powerful Spark/Hadoop clusters but requires more management and can be less cost-effective for nightly batch jobs compared to serverless Glue ETL.
- D. Glue DataBrew is primarily for visual data preparation and cleaning, not for full-scale automated ETL pipelines with Redshift loading.
AWS Glue ETL
A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development.
- Supports various data sources (S3, JDBC) and sinks (S3, Redshift).
- Automates schema inference and ETL job generation.
- Pay-as-you-go pricing based on DPU-hours.
Memory trick: Glue transforms batches and loads them into Redshift's house.