AWS Certified Solutions Architect – ProfessionalContinuously Improve Existing SolutionsMedium
A global manufacturing company has an existing data processing pipeline that relies on a self-managed Apache Spark cluster running on Amazon EC2 instances. The cluster requires constant monitoring, manual scaling adjustments, and patching, leading to high operational overhead and occasional resource bottlenecks during peak processing times. The company wants to improve the agility and cost-efficiency of its data processing by migrating to a fully managed, scalable, and serverless Spark solution on AWS. The solution must support existing Spark jobs without significant code changes. Which AWS service should a Solutions Architect recommend?
- AAmazon Redshift Spectrum
- BAWS Glue
- CAmazon Kinesis Data Analytics for Apache Flink
- DAmazon EMR on EC2
Show answer & explanationAnswer & explanation
Correct answer: B. AWS Glue
AWS Glue is a serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. It supports Apache Spark, allowing existing Spark jobs to run without significant code changes, and handles all provisioning, scaling, and patching, thus reducing operational overhead.
Why the other options are wrong
- A. Amazon Redshift Spectrum allows querying data in S3 using Redshift, but it's not a Spark-compatible processing engine for general-purpose data transformation jobs. It's for querying data lakes, not running Spark jobs.
- C. Amazon Kinesis Data Analytics for Apache Flink is for real-time stream processing using Apache Flink, not for batch processing with Apache Spark. It's a different technology and use case.
- D. Amazon EMR on EC2 still requires managing EC2 instances, including patching and some scaling considerations, which goes against the 'fully managed' and 'serverless' requirements to reduce operational overhead.
AWS Glue
A serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development, supporting Apache Spark.
- Fully managed, serverless Spark environment.
- Includes a Data Catalog for metadata management.
- Supports ETL (Extract, Transform, Load) operations.
- Pay-as-you-go pricing, scales automatically.
Memory trick: Glue 'sticks' your Spark jobs to a serverless cloud, no more server headaches!