AWS Certified Solutions Architect – Associate (SAA-C03)Design High-Performing ArchitecturesHard

A data engineering team is building a pipeline to ingest continuously flowing clickstream data from a high-traffic e-commerce website. The data needs to be captured, processed in real-time, and then loaded into a data lake for long-term analysis. The solution must be fully managed, scalable to handle petabytes of data, and provide built-in fault tolerance.

  1. AImplement Amazon Kinesis Data Firehose to stream data directly to Amazon S3.
  2. BUse Amazon SQS for ingestion and EC2 instances for processing.
  3. CStore data directly in Amazon Redshift and use Kinesis Data Analytics.
  4. DManually provision and manage Apache Kafka clusters on EC2 instances.
Show answer & explanation

Correct answer: A. Implement Amazon Kinesis Data Firehose to stream data directly to Amazon S3.

Amazon Kinesis Data Firehose is a fully managed service for delivering real-time streaming data to destinations like Amazon S3, Redshift, Splunk, and other custom HTTP endpoints. It automatically scales to match the throughput of your data, requires no administration, and batches/compresses data before delivery, making it ideal for efficiently loading petabytes of clickstream data into a data lake (S3) for long-term analysis.

Why the other options are wrong

  • B. SQS is a message queue, not optimized for continuous streaming data ingestion at petabyte scale, and managing EC2 instances for processing adds operational overhead.
  • C. Redshift is a data warehouse for analytical queries, not an ingestion service for raw clickstream data, and direct continuous loading can be inefficient. Kinesis Data Analytics is for processing, not primary ingestion/delivery to a data lake.
  • D. Manually managing Apache Kafka on EC2 instances involves significant operational overhead, which contradicts the 'fully managed' requirement.

Amazon Kinesis Data Firehose

Amazon Kinesis Data Firehose is a fully managed service that delivers real-time streaming data to destinations such as Amazon S3, Amazon Redshift, Splunk, and other custom HTTP endpoints.

  • Fully managed, no servers to manage.
  • Automatically scales to match data throughput.
  • Batches, compresses, and encrypts data before delivery.
  • Cost-effective for loading large volumes of streaming data into data lakes/warehouses.

Memory trick: Firehose: Fast, Fully-managed, Forwards data.

More Design High-Performing Architectures questions