AWS Certified Solutions Architect – ProfessionalContinuously Improve Existing SolutionsMedium

A global manufacturing company has an existing data processing pipeline that relies on a single, large Apache Spark cluster running on Amazon EC2 instances. This cluster is over-provisioned to handle peak loads, leading to high costs during off-peak times. The company wants to optimize costs and improve efficiency by dynamically scaling resources based on workload demands. The solution must support Spark workloads and allow for granular control over compute resources. Which architectural change should the solutions architect recommend?

  1. AMigrate the Spark workloads to AWS Glue for serverless ETL processing.
  2. BDeploy the Spark cluster on Amazon ECS with Fargate for serverless containers.
  3. CUtilize Amazon EMR with EC2 Spot Instances and managed scaling.
  4. DRe-implement the data processing logic using AWS Lambda functions.
Show answer & explanation

Correct answer: C. Utilize Amazon EMR with EC2 Spot Instances and managed scaling.

Amazon EMR is a managed service for running big data frameworks like Spark. EMR's managed scaling feature automatically adjusts the number of instances based on workload, and the use of EC2 Spot Instances can significantly reduce costs for fault-tolerant Spark jobs, directly addressing the cost optimization and dynamic scaling requirements.

Why the other options are wrong

  • A. AWS Glue is a serverless ETL service and can run Spark jobs, but it abstracts away many Spark-specific configurations and might not offer the 'granular control over compute resources' desired for an existing Spark workload, especially for advanced tuning. While cost-effective, it's a 'lift and shift' of the Spark environment.
  • B. Deploying Spark on ECS with Fargate introduces containerization and serverless compute, but managing a Spark cluster on ECS (even with Fargate) requires more manual effort for cluster orchestration, scaling, and integration compared to the fully managed and optimized scaling of Amazon EMR for Spark workloads.
  • D. Re-implementing Spark logic into AWS Lambda functions is a complete architectural change, likely a significant rewrite for complex Spark jobs, and Lambda's execution limits (15 minutes) are not suitable for long-running Spark tasks.

Amazon EMR for Spark Workloads

Amazon EMR is a managed cluster platform that simplifies running big data frameworks, such as Apache Spark and Hadoop, on AWS to process and analyze vast amounts of data.

  • Optimized for big data frameworks like Spark.
  • Offers managed scaling to adjust resources based on load.
  • Supports EC2 Spot Instances for cost optimization.
  • Reduces operational overhead compared to self-managed clusters.

Memory trick: EMR is the 'Master' of big data, scaling your Spark jobs on cost-saving Spot instances.

More Continuously Improve Existing Solutions questions