AWS Certified Developer – Associate (DVA-C02)Troubleshooting and MonitoringMedium
A developer is monitoring an application deployed on Amazon EC2 instances behind an Application Load Balancer (ALB). They observe an increasing number of HTTP 500 errors originating from the EC2 instances, as reported by the ALB's CloudWatch metrics. The EC2 instances show high CPU utilization and memory usage, but are not crashing. The application logs indicate frequent database connection timeouts. What action should the developer take FIRST to address these issues?
- AConfigure Auto Scaling to add more EC2 instances to distribute the load.
- BReview the database's CloudWatch metrics for connection limits, CPU, and I/O utilization.
- CImplement a retry mechanism with exponential backoff in the application for database calls.
- DIncrease the instance type of the EC2 instances to provide more CPU and memory.
Show answer & explanationAnswer & explanation
Correct answer: B. Review the database's CloudWatch metrics for connection limits, CPU, and I/O utilization.
Given that the application logs show frequent database connection timeouts while EC2 instances are under stress, the database is a strong candidate for the root cause. Reviewing its metrics will confirm if the database itself is overloaded or misconfigured, which is a critical first step before scaling application servers or adding retry logic.
Why the other options are wrong
- A. Scaling out EC2 instances would only send more requests to an already struggling database, potentially making the problem worse.
- C. Implementing retries can be helpful but won't solve an overloaded database; it might even exacerbate the problem by increasing load on an already struggling database.
- D. Increasing EC2 instance size might temporarily alleviate the symptom on the application server but won't solve the underlying database bottleneck.
Database Bottleneck
A database bottleneck occurs when the database cannot handle the volume or complexity of requests, leading to application slowdowns, errors, and connection timeouts.
- Often manifested as high CPU/memory/IOPS on the database.
- Application logs show connection errors or slow queries.
- Can be resolved by scaling the database, optimizing queries, or connection pooling.
Memory trick: When 5xxs strike, follow the request path, starting with the backend.