AWS Certified Developer – Associate (DVA-C02)Troubleshooting and MonitoringHard

A developer has implemented AWS CloudWatch alarms for an Amazon SQS queue. The alarms are configured to trigger when `NumberOfMessagesSent` exceeds a threshold, indicating high producer activity. However, during recent load tests, the SQS queue was flooded with messages, but the `NumberOfMessagesSent` alarm did not trigger. Other SQS metrics like `ApproximateNumberOfMessagesVisible` correctly show a large number of messages. What is the MOST probable reason the alarm did not trigger?

  1. AThe CloudWatch alarm is configured with an incorrect statistic, such as `Minimum` instead of `Sum`.
  2. BThe `NumberOfMessagesSent` metric represents messages successfully delivered to consumers, not messages sent to the queue.
  3. CThe messages were sent to a different SQS queue than the one the alarm is configured for.
  4. DThe CloudWatch alarm's evaluation period is too short, causing it to miss brief spikes.
Show answer & explanation

Correct answer: A. The CloudWatch alarm is configured with an incorrect statistic, such as `Minimum` instead of `Sum`.

To detect a high volume of messages sent to an SQS queue over a period, the `Sum` statistic should be used for the `NumberOfMessagesSent` metric. If the alarm is configured with `Average`, `Minimum`, or `Maximum` over a period, it might not correctly reflect the total volume and thus fail to trigger even with many messages sent.

Why the other options are wrong

  • B. This is incorrect. `NumberOfMessagesSent` explicitly measures the number of messages successfully added to the queue by producers, not messages delivered to consumers.
  • C. While possible, the question states that `ApproximateNumberOfMessagesVisible` correctly shows a large number of messages *in the queue*, implying the messages did reach the intended queue. The issue is with the alarm's configuration, not the target queue.
  • D. If the evaluation period was too short (e.g., 1 minute for a 5-minute spike), it might miss some events, but `NumberOfMessagesSent` accumulates, so it should still reflect activity over its period. A longer period might be needed for sustained load, but incorrect statistic is more fundamental.

CloudWatch Alarm Statistics

CloudWatch alarms evaluate metrics based on a chosen statistic (e.g., `Average`, `Sum`, `Minimum`, `Maximum`, `SampleCount`). Selecting the correct statistic is crucial for an alarm to accurately reflect the desired behavior of a metric and trigger reliably.

  • `Sum` is often used for count-based metrics (e.g., `NumberOfMessagesSent`, `Errors`).
  • `Average` is used for rate or duration metrics (e.g., `Latency`, `CPUUtilization`).
  • Incorrect statistic can lead to alarms not triggering or triggering falsely.

Memory trick: CloudWatch Alarm Silent: Check the Statistic, Period, and Threshold, or it's a fright.

More Troubleshooting and Monitoring questions