AWS Certified Machine Learning – SpecialtyData EngineeringHard

A financial institution needs to analyze customer transaction data to detect fraudulent activities. The data arrives from various online and offline channels in near real-time. The volume can spike significantly during peak hours. The solution requires processing these incoming transactions, enriching them with customer master data from a database, and then making the enriched data available for real-time fraud detection models. The solution must be serverless, highly scalable, and capable of handling varying throughput. Which AWS service combination would best meet these requirements?

  1. AAmazon Kinesis Data Streams -> AWS Lambda -> Amazon DynamoDB
  2. BAmazon Kinesis Data Firehose -> AWS Glue Streaming ETL -> Amazon S3
  3. CAmazon Kinesis Data Streams -> AWS Lambda -> Amazon Kinesis Data Analytics for Apache Flink
  4. DAmazon SQS -> AWS Batch -> Amazon Aurora
Show answer & explanation

Correct answer: C. Amazon Kinesis Data Streams -> AWS Lambda -> Amazon Kinesis Data Analytics for Apache Flink

Amazon Kinesis Data Streams provides the scalable, real-time ingestion layer for high-throughput data. AWS Lambda can be used to read from the stream, enrich transactions with customer master data (e.g., from DynamoDB or an in-memory cache), and then write back to another stream or pass to Kinesis Data Analytics. Amazon Kinesis Data Analytics for Apache Flink is ideal for real-time processing and running fraud detection models directly on streaming data with high scalability and low latency.

Why the other options are wrong

  • A. DynamoDB is a good storage for real-time lookups, but it's not the final destination for processed data for real-time models. This option lacks a dedicated real-time analytics component for complex fraud detection.
  • B. Kinesis Data Firehose is simpler for direct delivery to S3 but less suitable for complex real-time processing and enrichment with Lambda. Glue Streaming ETL is for near real-time, but KDA Flink is generally more performant for complex real-time analytics on streams.
  • D. Amazon SQS is a message queue, not optimized for high-throughput streaming. AWS Batch is for batch processing, not real-time. Amazon Aurora is a relational database, not typically used for direct real-time model serving on streaming data.

Real-time Stream Processing for ML

Building pipelines to ingest, process, and analyze high-volume, low-latency data streams in real-time for immediate machine learning inference or anomaly detection.

  • Ingestion: Kinesis Data Streams for high throughput.
  • Processing: AWS Lambda for lightweight transformations/enrichment.
  • Advanced Analytics: Kinesis Data Analytics for Apache Flink for complex real-time logic.
  • Low latency and high scalability are critical.

Memory trick: Streams flow, Lambda enriches, Flink detects the fraud that breaches.

More Data Engineering questions