AWS Certified Data Engineer – AssociateData Ingestion and TransformationHard
A data analytics team needs to build a robust and fault-tolerant pipeline to process customer interaction data from various social media platforms. The data arrives as semi-structured JSON, with varying schemas and nested structures. The processing involves schema inference, flattening nested data, enriching with internal customer IDs, and storing the transformed data in a data warehouse (Amazon Redshift). The pipeline must handle data volumes up to 10 TB daily and be able to recover from failures without data loss. Which AWS service is most appropriate for orchestrating and executing these complex transformations?
- AAmazon Kinesis Data Analytics for Apache Flink
- BAWS Step Functions with custom code
- CAWS Lambda with Python scripts
- DAWS Glue ETL
Show answer & explanationAnswer & explanation
Correct answer: D. AWS Glue ETL
AWS Glue ETL is a serverless data integration service that excels at processing large volumes of semi-structured data, performing schema inference, complex transformations like flattening and enrichment, and is highly fault-tolerant. It's built on Apache Spark, making it efficient for 10 TB daily volumes and integrates well with Redshift.
Why the other options are wrong
- A. Kinesis Data Analytics for Apache Flink is for real-time streaming data processing, not for batch processing of 10 TB daily data from social media platforms with varying schemas.
- B. AWS Step Functions orchestrates workflows, but the actual data transformation logic for 10 TB daily with complex schema handling is best executed by a dedicated ETL service like AWS Glue ETL, not custom code managed within Step Functions directly.
- C. AWS Lambda is suitable for smaller, short-lived tasks and would be inefficient and complex to manage for 10 TB daily processing with complex transformations and fault tolerance.
AWS Glue ETL
A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. It's built on Apache Spark.
- Serverless: no infrastructure to manage.
- Automatic schema discovery (Glue Crawler).
- Supports various data sources and targets, including complex transformations.
Memory trick: Glue makes messy data neat for the warehouse seat.