A manufacturing company wants to implement predictive maintenance for its industrial equipment. Sensors on the equipment generate telemetry data continuously, which needs to be ingested, processed in real-time to detect anomalies, and then stored for historical analysis and machine learning model training. The solution must be highly scalable, cost-effective, and provide insights with minimal operational overhead. Which architecture should a Solutions Architect recommend?
- AAWS IoT Greengrass for edge processing, Amazon SQS for queuing, Amazon EC2 for historical analysis, and Amazon S3 Intelligent-Tiering for storage.
- BAWS IoT Core for ingestion, Amazon Kinesis Data Streams for real-time processing, Amazon EC2 for anomaly detection, and Amazon RDS for storage.
- CAWS IoT Core for ingestion, AWS Lambda for real-time processing and anomaly detection, Amazon S3 for raw data storage, and Amazon Redshift for historical analysis.
- DAmazon Kinesis Data Firehose for ingestion, Amazon EMR for batch processing, Amazon DynamoDB for time-series data, and Amazon QuickSight for visualization.
Show answer & explanationAnswer & explanation
Correct answer: C. AWS IoT Core for ingestion, AWS Lambda for real-time processing and anomaly detection, Amazon S3 for raw data storage, and Amazon Redshift for historical analysis.
This architecture provides a serverless, scalable, and cost-effective solution. AWS IoT Core handles device connectivity and ingestion. AWS Lambda processes data in real-time for anomaly detection without managing servers. Amazon S3 offers highly scalable and cost-effective storage for raw data. Amazon Redshift is ideal for historical analysis and aggregates, supporting BI and ML training.
Why the other options are wrong
- A. IoT Greengrass is for edge processing, which is useful but doesn't replace the cloud architecture for central analysis. SQS is a message queue, not ideal for real-time stream processing and high-volume ingestion. EC2 for historical analysis adds operational overhead compared to Redshift. S3 Intelligent-Tiering is a storage class, not the primary storage solution for raw data lake.
- B. Kinesis Data Streams is suitable for real-time, but EC2 for anomaly detection adds operational overhead. RDS is not ideal for petabyte-scale time-series data and historical analysis due to cost and scalability limitations compared to S3 and Redshift.
- D. Kinesis Data Firehose is good for ingestion but primarily for delivery to destinations, not real-time processing itself. EMR is for batch processing, not real-time anomaly detection. DynamoDB can store time-series but S3/Redshift combination is more common for raw storage and historical analysis at scale. QuickSight is for visualization, not the core processing/storage.
IoT Predictive Maintenance Architecture
An AWS architecture for ingesting, processing, and analyzing IoT sensor data in real-time to predict equipment failures and optimize maintenance schedules.
- AWS IoT Core for secure device connectivity and data ingestion.
- AWS Lambda for serverless real-time data processing and anomaly detection.
- Amazon S3 for cost-effective, scalable raw data storage (data lake).
- Amazon Redshift for historical analysis, aggregation, and ML model training.
Memory trick: Connect with IoT, process with Lambda, store in S3, analyze with Redshift.