AWS Certified Machine Learning – SpecialtyData EngineeringHard
A data engineering team needs to transform raw streaming data from IoT devices into a structured format suitable for machine learning. The data arrives as JSON objects, but some fields are nested, and inconsistent schema versions occasionally appear. They need to flatten the nested structures, handle schema evolution, and convert the data into Apache Parquet format before storing it in Amazon S3. Which AWS service is best suited for performing these transformations in a serverless, scalable, and cost-effective manner?
- AAWS Lambda with custom Python scripts
- BAmazon EMR with Apache Spark Streaming
- CAWS Glue streaming ETL jobs
- DAmazon Kinesis Data Analytics Studio
Show answer & explanationAnswer & explanation
Correct answer: C. AWS Glue streaming ETL jobs
AWS Glue streaming ETL jobs are built on Apache Spark Streaming and are specifically designed for continuously processing streaming data. They provide schema inference, schema evolution handling, and powerful ETL capabilities to flatten nested JSON and convert data to formats like Parquet, all in a serverless and scalable manner, making them ideal for this scenario.
Why the other options are wrong
- A. While Lambda can process streams, handling complex schema evolution, nested JSON flattening, and Parquet conversion at scale with custom Python scripts can become complex and resource-intensive for Lambda's typical execution model.
- B. Amazon EMR with Apache Spark Streaming can perform these tasks, but it's not serverless and requires managing clusters, which adds operational overhead and potentially higher cost compared to AWS Glue's serverless model.
- D. Kinesis Data Analytics Studio (based on Apache Flink) can process streaming data and handle transformations, but AWS Glue's strong integration with the AWS Glue Data Catalog and its native support for complex ETL, including schema evolution and Parquet conversion, makes it a more comprehensive fit for this specific set of requirements.
AWS Glue Streaming ETL
A serverless data integration service that provides streaming ETL capabilities, built on Apache Spark Streaming, for transforming data from various sources to targets, handling schema evolution and complex data structures.
- Serverless Apache Spark Streaming environment.
- Handles schema inference and evolution.
- Ideal for flattening nested data and format conversions.
- Cost-effective and scalable for continuous data processing.
Memory trick: Glue streams and transforms, Spark-powered and serverless!