AWS Certified SysOps Administrator – AssociateMonitoring, Logging, and RemediationEasy

A company is running a critical application on Amazon EC2 instances that are part of an Auto Scaling group. They need to be notified immediately if any instance in the group fails its health checks and is marked as unhealthy by Auto Scaling. Which AWS service and configuration should a SysOps administrator use to meet this requirement with the lowest operational overhead?

  1. AConfigure Amazon CloudWatch Alarms to monitor the `CPUUtilization` metric for each instance and send notifications to an Amazon SNS topic.
  2. BConfigure an Amazon CloudWatch Alarm for the `EC2 Auto Scaling Group` namespace, monitoring the `UnhealthyHostCount` metric, and send notifications to an Amazon SNS topic.
  3. CImplement a custom script on each EC2 instance to periodically check its own health and send alerts via Amazon SES if issues are detected.
  4. DConfigure Amazon EventBridge to capture `EC2 Instance State-change Notification` events and filter for `terminated` status, then send to an Amazon SNS topic.
Show answer & explanation

Correct answer: B. Configure an Amazon CloudWatch Alarm for the `EC2 Auto Scaling Group` namespace, monitoring the `UnhealthyHostCount` metric, and send notifications to an Amazon SNS topic.

Monitoring the `UnhealthyHostCount` metric directly from the Auto Scaling group namespace is the most efficient way to detect and be notified about unhealthy instances. This metric specifically tracks instances that fail health checks within the group.

Why the other options are wrong

  • A. Monitoring CPUUtilization is a general health indicator but doesn't directly tell if Auto Scaling has marked an instance as unhealthy. It's also less efficient for group-wide health.
  • C. Implementing custom scripts on instances increases operational overhead, is less reliable, and duplicates functionality already provided by AWS services.
  • D. EventBridge can capture termination events, but an instance might be unhealthy and replaced by Auto Scaling without being explicitly 'terminated' in a way that triggers this specific event immediately as an 'unhealthy' state.

Auto Scaling Group Health Monitoring

Auto Scaling groups provide metrics to monitor the health of instances within the group, allowing for proactive notifications when instances become unhealthy.

  • The `UnhealthyHostCount` metric tracks instances failing health checks.
  • CloudWatch Alarms can be configured on Auto Scaling group metrics.
  • SNS topics are used for notifications from CloudWatch Alarms.

Memory trick: Auto Scaling's health is counted, CloudWatch alarms sound.

More Monitoring, Logging, and Remediation questions