A developer has implemented AWS CloudWatch alarms for an Amazon SQS queue. The alarms are configured to trigger when `NumberOfMessagesSent` exceeds a threshold, indicating high producer activity. However, during recent load tests, the SQS queue was flooded with messages, but the `NumberOfMessagesSent` alarm did not trigger. Other SQS metrics like `ApproximateNumberOfMessagesVisible` correctly show a large number of messages. What is the MOST probable reason the alarm did not trigger?
- AThe CloudWatch alarm is configured with an incorrect statistic, such as `Minimum` instead of `Sum`.
- BThe `NumberOfMessagesSent` metric represents messages successfully delivered to consumers, not messages sent to the queue.
- CThe messages were sent to a different SQS queue than the one the alarm is configured for.
- DThe CloudWatch alarm's evaluation period is too short, causing it to miss brief spikes.
Show answer & explanationAnswer & explanation
Correct answer: A. The CloudWatch alarm is configured with an incorrect statistic, such as `Minimum` instead of `Sum`.
To detect a high volume of messages sent to an SQS queue over a period, the `Sum` statistic should be used for the `NumberOfMessagesSent` metric. If the alarm is configured with `Average`, `Minimum`, or `Maximum` over a period, it might not correctly reflect the total volume and thus fail to trigger even with many messages sent.
Why the other options are wrong
- B. This is incorrect. `NumberOfMessagesSent` explicitly measures the number of messages successfully added to the queue by producers, not messages delivered to consumers.
- C. While possible, the question states that `ApproximateNumberOfMessagesVisible` correctly shows a large number of messages *in the queue*, implying the messages did reach the intended queue. The issue is with the alarm's configuration, not the target queue.
- D. If the evaluation period was too short (e.g., 1 minute for a 5-minute spike), it might miss some events, but `NumberOfMessagesSent` accumulates, so it should still reflect activity over its period. A longer period might be needed for sustained load, but incorrect statistic is more fundamental.
CloudWatch Alarm Statistics
CloudWatch alarms evaluate metrics based on a chosen statistic (e.g., `Average`, `Sum`, `Minimum`, `Maximum`, `SampleCount`). Selecting the correct statistic is crucial for an alarm to accurately reflect the desired behavior of a metric and trigger reliably.
- `Sum` is often used for count-based metrics (e.g., `NumberOfMessagesSent`, `Errors`).
- `Average` is used for rate or duration metrics (e.g., `Latency`, `CPUUtilization`).
- Incorrect statistic can lead to alarms not triggering or triggering falsely.
Memory trick: CloudWatch Alarm Silent: Check the Statistic, Period, and Threshold, or it's a fright.