A DevOps team manages several critical applications deployed on Amazon EC2 instances. They need to create a robust alerting system that notifies on-call engineers via SMS and PagerDuty when CPU utilization exceeds 90% for more than 5 minutes on any instance, or when a critical application log message (e.g., 'ERROR: Database connection failed') appears. The system must also automatically create a JIRA ticket for investigation. What is the most effective and integrated AWS solution?
- AImplement AWS Systems Manager Explorer and OpsCenter to aggregate operational data, then configure OpsCenter to create PagerDuty and JIRA tickets based on custom insights and CloudWatch Alarms.
- BUse CloudWatch Alarms for both CPU and log metrics, sending notifications to a single SNS topic. Configure SNS to directly send SMS, and use PagerDuty and JIRA webhooks subscribed to the SNS topic.
- CSet up CloudWatch Alarms for CPU utilization to trigger an SNS topic. For logs, create CloudWatch Subscription Filters to send critical log messages to a Lambda function, which then sends SMS, PagerDuty, and JIRA notifications.
- DConfigure CloudWatch Alarms for CPU utilization to trigger an SNS topic, and use a Lambda function to send SMS and create PagerDuty/JIRA tickets. For logs, use CloudWatch Log Metric Filters and another alarm.
Show answer & explanationAnswer & explanation
Correct answer: D. Configure CloudWatch Alarms for CPU utilization to trigger an SNS topic, and use a Lambda function to send SMS and create PagerDuty/JIRA tickets. For logs, use CloudWatch Log Metric Filters and another alarm.
This approach uses CloudWatch Alarms for both metric and log-based alerting, targeting SNS for fan-out. A single Lambda function can then process these SNS messages, parse the alert details, and integrate with external services like PagerDuty (via API) and JIRA (via API) for comprehensive notification and ticket creation. SNS can also directly send SMS. This centralizes the notification logic in Lambda, making it flexible and scalable.
Why the other options are wrong
- A. Systems Manager Explorer and OpsCenter are for aggregating operational data and managing operational items (OpsItems), but they don't natively provide the direct SMS/PagerDuty/JIRA notification and integration capabilities for arbitrary CloudWatch Alarms and log events without significant custom automation, which is typically handled by Lambda in this context.
- B. SNS can directly send SMS, but it cannot directly integrate with PagerDuty or JIRA via webhooks without an intermediate service like Lambda to format the payload correctly for their APIs.
- C. While CloudWatch Subscription Filters to Lambda is good for log processing, sending CPU alarms to a separate SNS topic and then duplicating SMS/PagerDuty/JIRA logic in potentially two Lambda functions (or making one Lambda handle both disparate sources) is less efficient than a single SNS topic fanning out to one Lambda for all complex integrations.
CloudWatch Log Metric Filters
Allows you to extract metric data from log events in CloudWatch Logs, enabling you to create CloudWatch Alarms based on patterns found in your logs.
- Transforms log data into numerical CloudWatch metrics.
- Enables alerting on specific log patterns (e.g., 'ERROR').
- Can count occurrences or extract values for metrics.
Memory trick: Alarms to SNS, Lambda Links Logs & Tickets.