AWS Certified Data Engineer – AssociateData Ingestion and TransformationHard
A marketing agency needs to process customer interaction data from various social media platforms. The data arrives in diverse formats (JSON, XML, CSV) and requires complex transformations, including data parsing, standardization, deduplication, and sentiment analysis using custom Python scripts. The processing needs to be flexible, allowing for custom code execution, and scalable to handle spikes in data volume. The transformed data should be stored in a data lake in Parquet format. Which AWS service provides the most flexibility and control for custom code-driven data transformation at scale?
- AAWS Glue ETL
- BAmazon Kinesis Data Analytics
- CAWS Lambda with Step Functions
- DAmazon EMR with Apache Spark
Show answer & explanationAnswer & explanation
Correct answer: D. Amazon EMR with Apache Spark
Amazon EMR with Apache Spark provides the most flexibility and control for custom code-driven data transformation. It allows running custom Python (PySpark) scripts for complex parsing, standardization, deduplication, and sentiment analysis at scale. While Glue ETL is serverless Spark, EMR offers more direct control over the cluster, libraries, and execution environment, which is crucial for highly custom and complex processing tasks.
Why the other options are wrong
- A. AWS Glue ETL is serverless Spark but might be less flexible for highly custom, dependency-heavy Python scripts compared to EMR where you control the cluster environment.
- B. Kinesis Data Analytics is for stream processing, primarily for SQL or Flink applications, not for batch processing with arbitrary custom Python scripts.
- C. Lambda is suitable for event-driven, short-running tasks. Complex, large-scale data transformations with custom Python scripts would exceed Lambda's limitations and be inefficient.
Amazon EMR with Apache Spark
A managed cluster platform that simplifies running big data frameworks, such as Apache Spark and Hadoop, on AWS to process vast amounts of data.
- Provides full control over cluster configuration and software.
- Supports various big data frameworks and languages (Spark, Hive, Presto, Python).
- Scales automatically to handle varying workloads.
Memory trick: EMR's Spark runs custom Python scripts on big data mountains.