A DevOps team manages a fleet of Amazon EC2 instances running a critical web application. They need to be notified immediately if any instance's CPU utilization exceeds 90% for five consecutive minutes. The notification should be sent to a specific Slack channel, and an automated action should be triggered to scale out the Auto Scaling Group. How should they configure this monitoring and alerting solution?
- ACreate a CloudWatch Alarm based on the 'CPUUtilization' metric, set the threshold to 90% for 5 consecutive minutes, link it to an SNS topic, and configure Slack integration via AWS Chatbot and an EC2 Auto Scaling policy.
- BEnable EC2 instance detailed monitoring, create a CloudWatch Logs metric filter for high CPU events, and use a CloudWatch Alarm to trigger SNS and a Lambda function for scaling.
- CDeploy a custom agent on each EC2 instance to monitor CPU, send alerts to an SQS queue, and then use a Lambda function to process the queue and send notifications/trigger scaling.
- DUse AWS Config rules to monitor CPU utilization, trigger a Lambda function for Slack notifications, and another Lambda for Auto Scaling Group actions.
Show answer & explanationAnswer & explanation
Correct answer: A. Create a CloudWatch Alarm based on the 'CPUUtilization' metric, set the threshold to 90% for 5 consecutive minutes, link it to an SNS topic, and configure Slack integration via AWS Chatbot and an EC2 Auto Scaling policy.
CloudWatch Alarms are the native way to monitor metrics like CPU utilization and trigger actions. Linking the alarm to an SNS topic allows for multiple subscribers, including AWS Chatbot for Slack notifications and an Auto Scaling policy for automated scaling actions, directly addressing all requirements efficiently.
Why the other options are wrong
- B. CloudWatch Logs metric filters are for analyzing log data, not directly for standard EC2 metrics like CPU utilization which are already available in CloudWatch metrics. Detailed monitoring is good, but the rest of the chain is inefficient.
- C. This is an overly complex and custom solution for a problem that CloudWatch Alarms and Auto Scaling policies solve natively and more efficiently.
- D. AWS Config is for compliance and configuration changes, not real-time metric-based alerting and scaling.
CloudWatch Alarm Actions
CloudWatch Alarms can trigger various actions when a metric goes into an 'ALARM' state, such as sending notifications or initiating automated responses.
- Can send notifications via SNS.
- Can trigger Auto Scaling actions (scale out/in).
- Can invoke EC2 actions (stop, terminate, reboot, recover).
- Can create Systems Manager OpsItems.
Memory trick: CloudWatch Alarm 'SNS-Chatbot-Scale' is the perfect trio for CPU spikes.