AWS Certified Solutions Architect – Associate (SAA-C03)Design Resilient ArchitecturesMedium

A company is building a new IoT platform that will ingest millions of small data points per second from connected devices globally. This data needs to be processed in real-time and then stored in a data lake for long-term analytics. The solution must be highly scalable, fault-tolerant, and capable of handling continuous high-throughput data ingestion. Which combination of AWS services should be used?

  1. AAWS IoT Core for ingestion, Amazon EC2 for processing, Amazon EBS for storage.
  2. BAmazon SQS for ingestion, AWS Lambda for processing, Amazon S3 for storage.
  3. CAmazon Kinesis Data Firehose for ingestion, AWS Glue for processing, Amazon Redshift for storage.
  4. DAmazon Kinesis Data Streams for ingestion, AWS Lambda for real-time processing, Amazon S3 for data lake storage.
Show answer & explanation

Correct answer: D. Amazon Kinesis Data Streams for ingestion, AWS Lambda for real-time processing, Amazon S3 for data lake storage.

Amazon Kinesis Data Streams is designed for real-time processing of large streams of data. AWS Lambda can process events from Kinesis Streams, providing serverless, scalable real-time processing. Amazon S3 is the ideal choice for a cost-effective, highly scalable, and durable data lake.

Why the other options are wrong

  • A. AWS IoT Core is for connecting devices, but EC2 and EBS are not the most scalable or cost-effective choices for processing and storage of millions of small data points per second in a real-time, serverless manner.
  • B. Amazon SQS is a message queuing service, suitable for individual messages but not optimized for high-throughput real-time streaming of millions of small data points per second.
  • C. Kinesis Data Firehose is for loading streaming data into data stores, not primarily for real-time processing. AWS Glue is for ETL, not real-time processing. Redshift is a data warehouse, not a data lake for raw, unstructured data.

Real-time Data Pipeline

A system designed to ingest, process, and store high-volume, continuous data streams with minimal latency, typically using services like Kinesis, Lambda, and S3.

  • Handles high-throughput data ingestion.
  • Processes data with low latency.
  • Stores raw or processed data for analytics.

Memory trick: Kinesis streams the data fast, Lambda processes, S3 makes it last.

More Design Resilient Architectures questions