AWS Certified Solutions Architect – ProfessionalContinuously Improve Existing SolutionsMedium

A global manufacturing company has an existing data processing pipeline that relies on a self-managed Apache Spark cluster running on-premises. The cluster is difficult to scale, costly to maintain, and experiences frequent resource contention, leading to delayed data analytics. The company wants to migrate this workload to AWS to improve scalability, reduce operational overhead, and accelerate data processing. Which AWS service should they use for their Spark workloads?

  1. AAmazon Redshift
  2. BAmazon EMR
  3. CAWS Lambda
  4. DAWS Glue
Show answer & explanation

Correct answer: B. Amazon EMR

Amazon EMR (Elastic MapReduce) is a fully managed service that makes it easy to run big data frameworks like Apache Spark, Hadoop, and Presto. It provides scalable, cost-effective, and highly available clusters, eliminating the operational overhead of self-managing Spark on-premises.

Why the other options are wrong

  • A. Amazon Redshift is a data warehouse service for analytical queries, not a platform for running Apache Spark processing jobs.
  • C. AWS Lambda is a serverless compute service for short-running functions, not suitable for long-running, distributed Apache Spark jobs.
  • D. AWS Glue is a serverless ETL service and can run Spark jobs, but EMR offers more direct control and optimization options for existing, complex Spark workloads and a wider range of big data frameworks.

Amazon EMR for Spark Workloads

Amazon EMR is a managed cluster platform that simplifies running big data frameworks like Apache Spark, Hadoop, and Presto on AWS.

  • Provides scalable and flexible clusters for big data processing.
  • Reduces operational overhead compared to self-managed clusters.
  • Supports various instance types and auto-scaling for cost optimization.

Memory trick: EMR is like a managed ranch for your Spark horses, letting them run free and fast.

More Continuously Improve Existing Solutions questions