AWS Certified Developer – Associate (DVA-C02)Troubleshooting and MonitoringMedium

A developer has configured an Amazon Kinesis Data Stream with 5 shards. A Lambda function processes records from this stream. Recently, the developer noticed an increasing trend in the 'IteratorAgeMilliseconds' metric in CloudWatch for the Lambda function, sometimes exceeding the configured Lambda timeout. This indicates that the Lambda function is falling behind in processing records. Which of the following is the MOST effective action to improve the processing rate and reduce iterator age?

  1. AIncrease the Lambda function's memory allocation to improve its processing speed.
  2. BIncrease the number of shards in the Kinesis Data Stream to allow for more parallel processing.
  3. CIncrease the 'Batch size' for the Lambda function's Kinesis trigger to process more records per invocation.
  4. DDecrease the 'Batch window' for the Lambda function's Kinesis trigger to invoke Lambda more frequently.
Show answer & explanation

Correct answer: B. Increase the number of shards in the Kinesis Data Stream to allow for more parallel processing.

An increasing 'IteratorAgeMilliseconds' indicates that the consumer (Lambda) is not keeping up with the producer. Kinesis Data Streams process records per shard. To improve the overall processing rate and allow for more concurrent Lambda invocations, you need to increase the number of shards. More shards mean Kinesis can deliver records to more concurrent Lambda instances, thus reducing the backlog.

Why the other options are wrong

  • A. While increasing memory *can* help if the Lambda is CPU/memory bound, the primary bottleneck when falling behind on Kinesis streams is often the number of parallel consumers, which is directly tied to the number of shards.
  • C. Increasing batch size can reduce the number of invocations, but if the processing itself is slow, it might exacerbate the problem by giving each Lambda more work, potentially increasing iterator age further if it can't keep up.
  • D. Decreasing the batch window would invoke Lambda more frequently, but if the underlying processing capacity (determined by shards) is the bottleneck, this might just lead to more throttled or failed invocations without actually increasing the processing rate.

Kinesis IteratorAgeMilliseconds

The CloudWatch metric 'IteratorAgeMilliseconds' for a Kinesis consumer (like Lambda) measures the age of the last record processed by the consumer, relative to its arrival time in the stream. An increasing trend indicates the consumer is falling behind the producer.

  • Measures lag between record arrival and processing.
  • High value means consumer is falling behind.
  • Affected by consumer capacity and stream throughput.

Memory trick: Kinesis needs more 'Shards' to speed up the 'Iterator' and catch the stream.

More Troubleshooting and Monitoring questions