AWS Certified DevOps Engineer – ProfessionalResilient Cloud SolutionsMedium

A global news organization operates a content management system (CMS) that serves dynamic web pages. The CMS backend is deployed on Amazon EC2 instances within an Auto Scaling group across multiple Availability Zones. During major breaking news events, the application experiences massive, unpredictable spikes in traffic, leading to degraded performance and occasional outages. The DevOps team needs to design a solution that can absorb these traffic spikes and automatically scale the backend quickly, ensuring high availability and consistent user experience. Which approach should they implement?

  1. AUtilize AWS Auto Scaling group with a target tracking scaling policy based on average CPU utilization.
  2. BIntegrate an Amazon CloudFront distribution with WAF rules to filter malicious traffic.
  3. CImplement an Amazon SQS queue to decouple the web tier from the processing tier.
  4. DIncrease the minimum and maximum capacity of the Auto Scaling group to a much higher fixed value.
Show answer & explanation

Correct answer: C. Implement an Amazon SQS queue to decouple the web tier from the processing tier.

Decoupling the web tier from the processing tier using an SQS queue allows the web servers to quickly offload requests into the queue, absorbing sudden traffic spikes. The backend processing instances can then pull messages from the queue at their own pace, scaling independently to handle the workload without overwhelming the web tier or the processing capacity.

Why the other options are wrong

  • A. While an Auto Scaling group with target tracking is good for scaling, it reacts to increased CPU utilization *on the instances*. If the traffic spike is so massive and sudden that it overwhelms the instances before they can scale (or if scaling takes too long), it won't effectively 'absorb' the initial impact and prevent degradation.
  • B. CloudFront (CDN) and WAF improve performance and security at the edge, but they don't directly address the backend's ability to process dynamic web pages or absorb massive, unpredictable spikes in processing demand at the application layer.
  • D. Increasing the fixed capacity might handle some spikes but will lead to over-provisioning and higher costs during normal periods, and it won't react dynamically to unpredictable spikes. It doesn't 'absorb' the spike, just prepares for it.

Decoupling with SQS

Decoupling application components using an Amazon SQS queue allows independent scaling, improves fault tolerance by buffering requests, and absorbs traffic spikes, preventing cascading failures.

  • SQS acts as a buffer between producers and consumers.
  • Producers can write messages quickly without waiting for consumers.
  • Consumers can process messages at their own rate, scaling independently.
  • Prevents system overload during traffic spikes and improves overall resilience.

Memory trick: SQS queue absorbs the wave, processes at its own pace.

More Resilient Cloud Solutions questions