AWS Certified Developer – Associate (DVA-C02)Troubleshooting and MonitoringHard

A developer is investigating an issue where an AWS Fargate service running a Docker container intermittently crashes and restarts. CloudWatch Container Insights metrics show that the service's memory utilization frequently spikes to 95% or more before the crashes, even though the container's `memoryReservation` is set to 512MB and `memoryLimit` to 1024MB. The application logs indicate no specific errors before the crash, only sudden termination. Which of the following is the MOST effective approach to diagnose and resolve this issue?

  1. AImplement a custom health check within the container to signal Fargate to restart the task proactively before a crash.
  2. BConfigure an Amazon CloudWatch alarm to notify when memory utilization exceeds 80% to detect issues earlier.
  3. CAnalyze the application's memory usage patterns within the container using profiling tools to identify memory leaks or excessive consumption.
  4. DIncrease the `memoryReservation` for the container to match `memoryLimit`, ensuring more memory is guaranteed.
Show answer & explanation

Correct answer: C. Analyze the application's memory usage patterns within the container using profiling tools to identify memory leaks or excessive consumption.

The symptoms (intermittent crashes, high memory utilization spikes before crash, no specific app errors) strongly suggest a memory leak or inefficient memory management within the application code itself. While increasing limits might temporarily hide the problem, it doesn't solve the root cause. Profiling the application's memory usage is the most effective way to diagnose and resolve the underlying issue.

Why the other options are wrong

  • A. Implementing a health check might restart the container slightly more gracefully, but it's a reactive measure. It doesn't prevent the memory issue from occurring or diagnose why it's happening, leading to continuous restarts and degraded service.
  • B. Setting an alarm is a good monitoring practice but only notifies about the symptom (high memory), it doesn't diagnose or resolve the root cause of the crash. The issue is already known.
  • D. Increasing `memoryReservation` to match `memoryLimit` would guarantee more memory, but if a leak exists, it would only delay the crash and potentially cost more, not solve the underlying memory leak or excessive consumption.

Container Memory Troubleshooting

Troubleshooting container crashes related to memory often involves distinguishing between insufficient allocated resources and inefficient application memory usage (e.g., memory leaks). Profiling tools are key for diagnosing application-specific memory issues.

  • High memory spikes followed by OOM (Out Of Memory) errors are indicative of application-level issues.
  • Container orchestration platforms (like Fargate/ECS) terminate tasks that exceed their memory limits.
  • Memory profiling helps identify leaks, inefficient data structures, or excessive allocations within the application.

Memory trick: Fargate Crash: Memory spikes hint at a leak; Profile to find the source and make it sleek.

More Troubleshooting and Monitoring questions