Professional Cloud ArchitectAnalyze and optimize technical and business processesMedium

A large e-commerce company is experiencing inconsistent application performance during peak shopping seasons. Their current architecture uses Compute Engine instances managed by a regional Managed Instance Group (MIG) behind a Global External Application Load Balancer. They observe that instances often reach 90% CPU utilization before new instances are provisioned, leading to customer-facing latency. Which optimization strategy should they implement to improve responsiveness during traffic spikes?

  1. ASwitch from a regional Managed Instance Group to a zonal Managed Instance Group.
  2. BConfigure the autoscaling policy to scale based on HTTP load balancing serving capacity.
  3. CImplement a custom autoscaling metric based on application-level queue length.
  4. DIncrease the minimum number of instances in the Managed Instance Group.
Show answer & explanation

Correct answer: B. Configure the autoscaling policy to scale based on HTTP load balancing serving capacity.

To improve responsiveness during traffic spikes, configuring autoscaling based on HTTP load balancing serving capacity allows the MIG to provision new instances proactively, before CPU utilization becomes a bottleneck. This metric considers the load balancer's view of traffic and can scale out more quickly than reactive CPU-based scaling.

Why the other options are wrong

  • A. Switching to a zonal MIG would reduce redundancy and disaster recovery capabilities, and does not directly address the autoscaling responsiveness issue.
  • C. Implementing a custom autoscaling metric based on application-level queue length is a valid advanced strategy but is more complex and might not be as immediately effective as leveraging the built-in HTTP load balancing metric for this specific scenario.
  • D. Increasing the minimum number of instances helps with baseline capacity but does not address reactive scaling during unpredictable spikes.

Autoscaling Metrics

Autoscaling policies define how Managed Instance Groups (MIGs) adjust the number of instances based on demand, using various metrics to trigger scaling actions.

  • CPU utilization is a common, but reactive, metric.
  • HTTP load balancing serving capacity is a proactive metric for web applications.
  • Custom metrics can provide fine-grained control based on application-specific needs.

Memory trick: React fast, scale smart, keep apps tart!

More Analyze and optimize technical and business processes questions