A large manufacturing company wants to implement a new IoT solution for factory automation. The solution must collect telemetry data from thousands of sensors (e.g., temperature, pressure, vibration) at a high frequency (tens of thousands of messages per second). This data needs to be processed in real-time for anomaly detection and machine learning model inference, and then stored for long-term historical analysis. The solution must be highly scalable, fault-tolerant, and cost-effective. Which AWS services should be used to build this data ingestion and processing pipeline?
- AAWS IoT Core for device connectivity, Amazon Kinesis Data Streams for ingestion, AWS Lambda for real-time processing, and Amazon S3 for long-term storage.
- BAWS IoT Core for device connectivity, Amazon MSK for ingestion, AWS Glue for real-time processing, and Amazon Redshift for long-term storage.
- CAWS IoT Core for device connectivity, Amazon Kinesis Data Firehose for ingestion, and Amazon S3 for storage.
- DAWS IoT Greengrass for edge processing, Amazon SQS for message queuing, and Amazon DynamoDB for all data storage.
Show answer & explanationAnswer & explanation
Correct answer: A. AWS IoT Core for device connectivity, Amazon Kinesis Data Streams for ingestion, AWS Lambda for real-time processing, and Amazon S3 for long-term storage.
AWS IoT Core handles device connectivity and message routing. Amazon Kinesis Data Streams is ideal for ingesting high-volume, real-time streaming data with fault tolerance. AWS Lambda can process this data in real-time for anomaly detection and ML inference. Finally, Amazon S3 provides highly scalable and cost-effective long-term storage for historical analysis. This combination addresses all requirements for scalability, real-time processing, and cost-effectiveness.
Why the other options are wrong
- B. Amazon MSK (Managed Streaming for Apache Kafka) can be used for ingestion but might be overkill or more complex than Kinesis Data Streams for this scenario. AWS Glue is typically for batch ETL, not real-time processing of individual messages. Redshift is a data warehouse, potentially more expensive for raw historical IoT data storage than S3.
- C. Kinesis Data Firehose is primarily for loading streaming data into data stores, not for real-time processing of individual messages for anomaly detection. It's better for batch delivery.
- D. IoT Greengrass is for edge computing, not the primary cloud ingestion. Amazon SQS is a message queue, not optimized for high-throughput streaming data ingestion. DynamoDB is not ideal for cost-effective long-term historical analysis of large volumes of time-series data.
IoT Data Ingestion and Processing Pipeline
An architecture designed to collect, process, and store high volumes of real-time data from IoT devices using serverless and managed AWS services.
- Uses IoT Core for device connectivity.
- Leverages Kinesis for high-throughput streaming ingestion.
- Employs Lambda for real-time event-driven processing.
Memory trick: IoT Core connects, Kinesis streams, Lambda processes, S3 stores.