AWS Certified Solutions Architect – ProfessionalDesign for New SolutionsHard

A global manufacturing company needs to implement a new data platform to collect, process, and analyze sensor data from its machinery across multiple factories worldwide. The data volume is expected to be in the tens of terabytes daily, requiring real-time ingestion, transformation, and storage for both immediate operational insights and long-term historical analysis. The solution must support SQL-based querying for analysts and provide a cost-effective storage solution for cold data. What is the MOST appropriate data platform architecture?

  1. AAWS IoT Core for data ingestion, Amazon Kinesis Data Streams for real-time processing with AWS Lambda, Amazon Timestream for hot data, and Amazon S3 with Athena for cold data analysis.
  2. BAWS IoT Core for data ingestion, Amazon Kinesis Data Firehose for streaming to Amazon S3, AWS Glue for ETL, and Amazon Redshift for analytical queries.
  3. CAWS IoT Core for data ingestion, Amazon MSK for streaming, Amazon EMR for processing, and Amazon Aurora PostgreSQL for all data storage.
  4. DAWS IoT Greengrass for edge processing, Amazon SQS for message queuing, Amazon DynamoDB for hot data, and Amazon Glacier for cold data.
Show answer & explanation

Correct answer: A. AWS IoT Core for data ingestion, Amazon Kinesis Data Streams for real-time processing with AWS Lambda, Amazon Timestream for hot data, and Amazon S3 with Athena for cold data analysis.

This architecture provides a robust solution. AWS IoT Core handles device connectivity and message ingestion. Kinesis Data Streams is ideal for high-volume, real-time streaming data ingestion, and AWS Lambda can process this data instantly for operational insights. Amazon Timestream is a purpose-built time-series database optimized for high-volume, low-latency ingestion and querying of time-series data, making it excellent for 'hot' operational data. For 'cold' long-term storage and cost-effective SQL-based analysis, Amazon S3 combined with Amazon Athena is a perfect fit, allowing analysts to query historical data directly from S3 without loading it into a data warehouse.

Why the other options are wrong

  • B. Kinesis Data Firehose is better for batch loading, not real-time processing of individual events, and Redshift can be expensive for petabyte-scale raw sensor data compared to S3/Athena.
  • C. MSK and EMR are powerful but can be more complex and costly than Kinesis/Lambda for this scenario. Aurora PostgreSQL is a relational database not optimized for time-series data at this scale, nor for cost-effective cold storage.
  • D. IoT Greengrass is for edge processing, not the core cloud ingestion. SQS is a queue, not a streaming service for tens of terabytes daily. DynamoDB is not optimized for time-series data, and Glacier is too slow for analytical queries.

IoT Time-Series Data Platform

An architecture for ingesting, processing, and analyzing high-volume, real-time time-series data from IoT devices, often utilizing specialized databases and tiered storage.

  • Leverages IoT Core and Kinesis for ingestion.
  • Uses Timestream for hot time-series data.
  • Combines S3 and Athena for cost-effective cold data analysis.

Memory trick: IoT Core streams to Kinesis, Lambda processes, Timestream holds hot, S3/Athena cools.

More Design for New Solutions questions