AWS Certified Developer – Associate (DVA-C02)Troubleshooting and MonitoringHard

A developer has implemented AWS CloudWatch alarms for an Amazon SQS queue. The alarms are configured to trigger when the 'ApproximateNumberOfMessagesVisible' metric exceeds 100 for 5 consecutive periods of 1 minute. However, during a recent incident, the queue backlog grew significantly (over 1000 messages), but the CloudWatch alarm did not trigger. Upon investigation, the developer found that the 'ApproximateNumberOfMessagesVisible' metric never consistently stayed above 100 for the full 5-minute period, even though the total messages in the queue were high. Which of the following is the MOST likely reason the alarm did not trigger?

  1. AThe Lambda function consuming messages from the queue was actively processing messages, reducing the 'visible' count.
  2. BThe CloudWatch alarm's 'Datapoints to alarm' setting was configured incorrectly, requiring more data points.
  3. CThe SQS queue was configured as a FIFO queue, which has different metric emission behavior.
  4. DThe 'Period' for the CloudWatch alarm is too long, causing it to miss transient spikes in visible messages.
Show answer & explanation

Correct answer: A. The Lambda function consuming messages from the queue was actively processing messages, reducing the 'visible' count.

The 'ApproximateNumberOfMessagesVisible' metric represents messages that are available for immediate consumption. If a consumer (e.g., a Lambda function) is actively processing messages, those messages become 'in flight' (hidden by the visibility timeout) and are no longer counted as 'visible'. Therefore, even with a large backlog, if the consumer is working efficiently, the number of *visible* messages might not consistently stay above the alarm threshold, leading to the alarm not triggering.

Why the other options are wrong

  • B. The question states it was '5 consecutive periods of 1 minute', implying 'Datapoints to alarm' was 5, which is standard. The issue is the metric's value, not the alarm's threshold logic.
  • C. FIFO queues have 'ApproximateNumberOfMessagesVisible' as well, and their behavior regarding visibility timeout is similar to standard queues in this context.
  • D. The 'Period' is 1 minute, which is standard. If anything, a shorter period might catch more transient spikes, but the core issue is the metric itself not reflecting the *total* backlog due to active consumption.

CloudWatch Alarm Statistics (SQS)

CloudWatch alarms based on SQS metrics like 'ApproximateNumberOfMessagesVisible' can behave counter-intuitively if consumers are active. 'ApproximateNumberOfMessagesVisible' counts messages ready for consumption, not messages currently 'in flight' or being processed. An active consumer can keep this metric low even with a large total queue backlog.

  • 'Visible' means available to be received.
  • Messages 'in flight' are not 'visible'.
  • Active consumers reduce 'visible' count.
  • Consider 'ApproximateNumberOfMessagesNotVisible' or 'ApproximateNumberOfMessagesDelayed' for total backlog.

Memory trick: SQS 'Visible' messages are like 'Hide-and-Seek'; if the consumer is good, they stay hidden.

More Troubleshooting and Monitoring questions