AWS Certified Data Engineer – AssociateData Operations and MonitoringHard

A data analytics company processes large volumes of streaming data using Amazon Kinesis Data Streams and then delivers this data to Amazon S3 for archival and further processing. They have strict requirements for data integrity and ensuring that no data is lost or duplicated during the transfer from Kinesis to S3. They use an AWS Lambda function to consume records from Kinesis and write them to S3. Occasionally, due to network issues or S3 API throttling, the Lambda function fails to deliver a batch of records to S3. The team needs to implement a solution to prevent data loss and ensure exactly-once delivery semantics where possible, and at-least-once otherwise, for records from Kinesis to S3, with minimal custom code.

  1. AConfigure the Lambda function to write data to an SQS queue first, then another Lambda to process from SQS to S3.
  2. BConfigure the Kinesis event source mapping for the Lambda function with maximum retry attempts and a destination for failed records, and implement idempotency in the Lambda.
  3. CImplement a custom retry mechanism with exponential backoff and a Dead-Letter Queue (DLQ) within the existing Lambda function.
  4. DUtilize Amazon Kinesis Data Firehose to directly deliver data from Kinesis Data Streams to S3.
Show answer & explanation

Correct answer: B. Configure the Kinesis event source mapping for the Lambda function with maximum retry attempts and a destination for failed records, and implement idempotency in the Lambda.

For Kinesis-to-Lambda-to-S3, the Kinesis event source mapping handles retries for failed Lambda invocations. Combining this with a destination for failed records (like SQS or SNS) prevents data loss. To address potential duplicates from retries and achieve near exactly-once semantics, the Lambda function should be idempotent, meaning it can process the same record multiple times without adverse effects (e.g., by using record IDs for conditional writes to S3).

Why the other options are wrong

  • A. Adding SQS adds another layer of complexity and potential for duplication if not handled carefully, and doesn't inherently solve exactly-once delivery to S3 from Kinesis.
  • C. While custom retry and DLQ in Lambda are good, the Kinesis event source mapping provides more robust, stream-aware retry capabilities, and the core issue of potential duplicates (from retries) needs idempotency, which isn't covered by just retries/DLQ.
  • D. Kinesis Data Firehose is an excellent service for Kinesis-to-S3 delivery, but the scenario explicitly mentions using a Lambda function for processing. Switching to Firehose would be a re-architecture, not an enhancement of the existing Lambda-based pipeline.

Kinesis-Lambda-S3 Reliability

Ensure reliable Kinesis-to-Lambda-to-S3 delivery by configuring Kinesis event source mapping retries/failure destinations and implementing idempotency in the Lambda function to handle potential duplicates.

  • ESM retries prevent data loss from transient Lambda failures.
  • Failure destinations capture unprocessable records.
  • Idempotency in Lambda handles duplicate processing from retries.

Memory trick: ESM retries, Lambda's idempotent key, keeps the data safe and true.

More Data Operations and Monitoring questions