A SysOps team manages a critical application running on multiple Amazon EC2 instances behind an Application Load Balancer (ALB). They observe that during peak traffic, users report slow response times, but CPU utilization on EC2 instances remains low. The team needs to identify the bottleneck. Which monitoring metric and service should they investigate first?
- AEC2 'NetworkIn' and 'NetworkOut' metrics using Amazon CloudWatch.
- BEBS 'BurstBalance' and 'VolumeQueueLength' metrics using Amazon CloudWatch.
- CALB 'TargetConnectionErrorCount' and 'HealthyHostCount' metrics using Amazon CloudWatch.
- DRDS 'DatabaseConnections' and 'CPUUtilization' metrics using Amazon CloudWatch.
Show answer & explanationAnswer & explanation
Correct answer: C. ALB 'TargetConnectionErrorCount' and 'HealthyHostCount' metrics using Amazon CloudWatch.
Since CPU utilization is low but users report slow response times, the bottleneck might be at the Application Load Balancer (ALB) or the connection to the target instances. 'TargetConnectionErrorCount' indicates issues establishing connections with target instances, and 'HealthyHostCount' shows if enough healthy instances are available to serve requests. These metrics directly point to potential issues with the ALB's ability to forward requests or the targets' ability to accept them, which aligns with slow response times despite low EC2 CPU.
Why the other options are wrong
- A. NetworkIn/NetworkOut are general network metrics for EC2, but slow response times with low CPU often point to issues *before* or *at* the instance, not necessarily the instance's overall network throughput.
- B. EBS metrics are relevant for disk I/O bottlenecks. While possible, slow response times with low CPU often suggest network/connection issues before storage issues, especially with an ALB in front.
- D. RDS metrics are relevant if the database is the bottleneck, but the scenario focuses on web application response times and low EC2 CPU, suggesting a front-end or connection issue first.
ALB Bottleneck Metrics
Key Amazon CloudWatch metrics for Application Load Balancers (ALB) to identify issues related to target connectivity, health, and request processing that can cause slow application response times.
- 'TargetConnectionErrorCount' identifies backend connection failures.
- 'HealthyHostCount' shows available healthy targets.
- These help diagnose issues when EC2 CPU is low but performance is poor.
Memory trick: ALB's health and connection errors reveal the hidden slowness.