AWS Certified Solutions Architect – Associate (SAA-C03)Design Resilient ArchitecturesMedium
A company is building a new IoT platform that will ingest millions of small data points per second from connected devices globally. This data needs to be processed in real-time and then stored in a data lake for long-term analytics. The solution must be highly scalable, fault-tolerant, and capable of handling continuous high-throughput data ingestion. Which combination of AWS services should be used?
- AAWS IoT Core for ingestion, Amazon EC2 for processing, Amazon EBS for storage.
- BAmazon SQS for ingestion, AWS Lambda for processing, Amazon S3 for storage.
- CAmazon Kinesis Data Firehose for ingestion, AWS Glue for processing, Amazon Redshift for storage.
- DAmazon Kinesis Data Streams for ingestion, AWS Lambda for real-time processing, Amazon S3 for data lake storage.
Show answer & explanationAnswer & explanation
Correct answer: D. Amazon Kinesis Data Streams for ingestion, AWS Lambda for real-time processing, Amazon S3 for data lake storage.
Amazon Kinesis Data Streams is designed for real-time processing of large streams of data. AWS Lambda can process events from Kinesis Streams, providing serverless, scalable real-time processing. Amazon S3 is the ideal choice for a cost-effective, highly scalable, and durable data lake.
Why the other options are wrong
- A. AWS IoT Core is for connecting devices, but EC2 and EBS are not the most scalable or cost-effective choices for processing and storage of millions of small data points per second in a real-time, serverless manner.
- B. Amazon SQS is a message queuing service, suitable for individual messages but not optimized for high-throughput real-time streaming of millions of small data points per second.
- C. Kinesis Data Firehose is for loading streaming data into data stores, not primarily for real-time processing. AWS Glue is for ETL, not real-time processing. Redshift is a data warehouse, not a data lake for raw, unstructured data.
Real-time Data Pipeline
A system designed to ingest, process, and store high-volume, continuous data streams with minimal latency, typically using services like Kinesis, Lambda, and S3.
- Handles high-throughput data ingestion.
- Processes data with low latency.
- Stores raw or processed data for analytics.
Memory trick: Kinesis streams the data fast, Lambda processes, S3 makes it last.