AWS Certified Solutions Architect – ProfessionalDesign Solutions for Organizational ComplexityHard

A financial services company is developing a new analytics platform on AWS. The platform needs to ingest large volumes of real-time market data from various external sources, process it for fraud detection, and store it for historical analysis. The data ingestion must be highly available and support varying throughput, while the processing logic needs to be executed reliably and handle potential failures without data loss. The historical data storage must be cost-effective and allow for complex queries. Which combination of AWS services should be used to build this analytics platform?

  1. AAWS DataSync for ingestion, AWS Batch for processing, and Amazon EBS for historical storage.
  2. BAmazon Kinesis Data Streams for ingestion, AWS Lambda for processing, and Amazon S3 with Amazon Athena for historical storage and queries.
  3. CAmazon SQS for ingestion, Amazon EC2 for processing, and Amazon RDS for historical storage and queries.
  4. DAmazon API Gateway for ingestion, Amazon EMR for processing, and Amazon Redshift for historical storage and queries.
Show answer & explanation

Correct answer: B. Amazon Kinesis Data Streams for ingestion, AWS Lambda for processing, and Amazon S3 with Amazon Athena for historical storage and queries.

Amazon Kinesis Data Streams is ideal for highly available, real-time data ingestion with varying throughput. AWS Lambda provides scalable, event-driven processing for fraud detection. Amazon S3 offers cost-effective, durable storage for historical data, and Amazon Athena enables complex queries directly on S3 data without managing servers.

Why the other options are wrong

  • A. AWS DataSync is for large-scale data transfer to S3, not real-time streaming ingestion. AWS Batch is for batch processing, not real-time, and EBS is block storage, not suitable for cost-effective historical data lakes.
  • C. SQS is a message queue, not optimized for streaming large volumes of real-time data. RDS is not cost-effective for large-scale historical data analysis compared to S3/Athena.
  • D. API Gateway is for API management, not high-volume streaming ingestion. Amazon EMR is for big data processing, but Lambda is more suitable for real-time, event-driven fraud detection. Amazon Redshift can be used for data warehousing, but S3/Athena is more cost-effective for a 'data lake' approach for historical analysis.

Real-time Analytics Pipeline

A real-time analytics pipeline on AWS typically involves services for high-throughput data ingestion (e.g., Kinesis), serverless or stream-processing compute (e.g., Lambda, Kinesis Analytics), and cost-effective, scalable storage for historical data (e.g., S3) with query capabilities (e.g., Athena).

  • Handles high-volume, high-velocity data
  • Enables immediate processing and insights
  • Decouples components for scalability and resilience
  • Leverages managed services for operational efficiency

Memory trick: Kinesis, Lambda, S3, Athena: The K.L.S.A. pipeline.

More Design Solutions for Organizational Complexity questions