AWS Certified DevOps Engineer – ProfessionalResilient Cloud SolutionsMedium
A global e-commerce company uses Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB) for its product catalog service. During peak sales events, the service experiences intermittent latency spikes and occasional errors, even though CPU utilization on the EC2 instances remains low. Further investigation reveals that the latency spikes correlate with a rapid increase in traffic, and the Auto Scaling group often takes several minutes to launch new instances and register them with the ALB. Which configuration adjustment would most effectively mitigate these latency spikes and improve responsiveness during traffic surges?
- AEnable `Predictive Scaling` for the Auto Scaling group based on historical traffic patterns.
- BIncrease the `min` capacity of the Auto Scaling group to handle expected peak traffic.
- CImplement a `Step Scaling Policy` that adds a fixed number of instances when CPU utilization exceeds 50%.
- DConfigure a `Target Tracking Scaling Policy` based on ALB `RequestCountPerTarget` metric.
Show answer & explanationAnswer & explanation
Correct answer: A. Enable `Predictive Scaling` for the Auto Scaling group based on historical traffic patterns.
The problem describes latency spikes during rapid traffic increases and slow Auto Scaling group response. Predictive Scaling uses machine learning to forecast future traffic and proactively scale capacity, thereby reducing reactive scaling delays.
Why the other options are wrong
- B. Increasing `min` capacity helps with baseline traffic but doesn't dynamically adjust for unexpected surges or prevent latency spikes caused by sudden, high-volume traffic increases.
- C. Step scaling based on CPU utilization is reactive and might be too slow for rapid traffic surges, as CPU might be low while response time is high due to connection backlog or other non-CPU bottlenecks.
- D. Target Tracking on `RequestCountPerTarget` is a good reactive scaling policy but still needs time to provision new instances, which is the core issue identified.
Predictive Scaling
An Auto Scaling policy that uses machine learning to forecast future traffic and scale capacity proactively, anticipating demand changes.
- Leverages historical data for forecasting.
- Launches instances before traffic spikes occur.
- Reduces latency and improves application availability during anticipated surges.
Memory trick: Predictive Powers Prevent Performance Plunges.