AWS Certified Solutions Architect – ProfessionalDesign for New SolutionsHard

A financial institution is building a new data platform to ingest, process, and analyze vast amounts of financial transaction data. The data arrives continuously from various sources, requiring real-time processing and analysis. The platform must support petabyte-scale storage, complex SQL queries for historical data, and machine learning workloads. Data governance and auditing are critical, and data must be encrypted at rest and in transit. Which combination of AWS services provides the most suitable architecture?

  1. AAmazon Kinesis for ingestion, Amazon SQS for queuing, Amazon EC2 for processing, and Amazon RDS for analytics.
  2. BAWS DataSync for ingestion, Amazon EMR for processing, Amazon Redshift for analytics, and Amazon S3 for cold storage.
  3. CAWS Database Migration Service (DMS) for ingestion, AWS Lambda for processing, Amazon DynamoDB for data storage, and Amazon QuickSight for analytics.
  4. DAmazon Kinesis Data Firehose for ingestion, AWS Glue for ETL, Amazon S3 for raw data lake, Amazon Redshift for data warehousing, and Amazon SageMaker for ML.
Show answer & explanation

Correct answer: D. Amazon Kinesis Data Firehose for ingestion, AWS Glue for ETL, Amazon S3 for raw data lake, Amazon Redshift for data warehousing, and Amazon SageMaker for ML.

This architecture provides a comprehensive, scalable, and secure solution. Kinesis Data Firehose efficiently ingests real-time streaming data. S3 acts as a scalable data lake for raw and processed data. AWS Glue provides serverless ETL for data transformation. Amazon Redshift is a petabyte-scale data warehouse optimized for complex SQL queries and analytics. SageMaker supports machine learning workloads. All services support encryption and integrate with AWS security and governance tools.

Why the other options are wrong

  • A. SQS is a message queue, not ideal for real-time streaming ingestion. EC2 for processing requires significant operational overhead. RDS is not designed for petabyte-scale data warehousing and complex analytical queries.
  • B. DataSync is primarily for large-scale data transfer to S3, not continuous real-time ingestion from diverse sources. EMR can process data but AWS Glue is often preferred for serverless ETL. Redshift is suitable, but S3 as cold storage only misses its role as the data lake for raw data.
  • C. DMS is for database migrations, not general real-time ingestion from diverse sources. Lambda is suitable for event-driven processing but may not handle petabyte-scale continuous data processing efficiently without complex orchestration. DynamoDB is a NoSQL database, not optimized for complex SQL analytics over historical data. QuickSight is a BI tool, not the core data storage/processing platform.

Real-time Data Lake Architecture

An AWS architectural pattern for ingesting, processing, and analyzing large volumes of streaming data in real-time, combining a data lake with a data warehouse and ML capabilities.

  • Uses Kinesis for real-time ingestion.
  • S3 for scalable, cost-effective data lake storage.
  • AWS Glue for serverless ETL and cataloging.
  • Redshift for analytical queries over structured data.
  • SageMaker for machine learning workloads.

Memory trick: Stream with Kinesis, lake with S3, transform with Glue, warehouse with Redshift, learn with SageMaker.

More Design for New Solutions questions