AWS Certified DevOps Engineer – ProfessionalIncident and Event ResponseMedium

A DevOps team manages a critical microservices application deployed on Amazon ECS. They need to ensure that if a container experiences a high number of restarts or crashes, an automated remediation process is triggered to replace the unhealthy container and notify the operations team. Which approach should they implement to achieve this?

  1. AUtilize ECS service auto-scaling policies to scale down the service when container health checks fail.
  2. BConfigure a CloudWatch Alarm on the ECS service's CPUUtilization metric, triggering an SNS topic for notification.
  3. CImplement an Amazon EventBridge rule that detects ECS service action events (e.g., 'ECS Task State Change') for container restarts/crashes, triggering an AWS Lambda function to update the service and an SNS topic for notification.
  4. DSet up AWS Config rules to monitor ECS task definitions for changes, and use AWS Systems Manager Automation to revert to a previous task definition if issues are detected.
Show answer & explanation

Correct answer: C. Implement an Amazon EventBridge rule that detects ECS service action events (e.g., 'ECS Task State Change') for container restarts/crashes, triggering an AWS Lambda function to update the service and an SNS topic for notification.

EventBridge can detect specific ECS events, such as task state changes indicating restarts or crashes. This can then trigger a Lambda function to perform remediation (e.g., forcing a new deployment to replace unhealthy containers) and send notifications via SNS, creating an automated and targeted response.

Why the other options are wrong

  • A. ECS service auto-scaling primarily scales tasks based on metrics like CPU or memory. While it can react to health checks, scaling DOWN when containers are unhealthy might worsen availability, and it doesn't explicitly trigger a replacement and notification for specific container issues.
  • B. CPUUtilization is a general performance metric and doesn't directly indicate container restarts/crashes. This approach would not address the specific problem of unhealthy containers.
  • D. AWS Config monitors resource configurations and changes, not real-time container health or crash events. While useful for compliance, it's not designed for automated, real-time remediation of unhealthy running containers.

Event-Driven Container Remediation

Using AWS EventBridge to detect specific container events (like restarts or crashes) and trigger automated actions via AWS Lambda to remediate the issue and notify relevant teams.

  • EventBridge monitors ECS Task State Change events.
  • Lambda functions execute custom remediation logic.
  • SNS provides immediate notification to operations teams.
  • Ensures application health and availability in containerized environments.

Memory trick: EventBridge sees a crash, Lambda fixes the task, SNS sends the alert.

More Incident and Event Response questions