AWS Certified Data Engineer – AssociateData Storage and ManagementEasy

A data engineering team is designing a data lake on AWS. They need to store streaming data from various IoT devices, which arrives at a high velocity and requires immediate processing for real-time analytics. The solution must be highly scalable and durable. Which AWS service is most appropriate for ingesting this data?

  1. AAmazon DynamoDB
  2. BAmazon S3
  3. CAmazon Redshift
  4. DAmazon Kinesis Data Firehose
Show answer & explanation

Correct answer: D. Amazon Kinesis Data Firehose

Amazon Kinesis Data Firehose is specifically designed for ingesting and loading streaming data into data lakes, data stores, and analytics services. It handles high velocity data and automatically scales to meet throughput demands without requiring ongoing administration.

Why the other options are wrong

  • A. Amazon DynamoDB is a NoSQL database for transactional data, not primarily for general-purpose streaming ingestion into a data lake.
  • B. Amazon S3 is suitable for static data storage, not direct ingestion of high-velocity streaming data for immediate processing.
  • C. Amazon Redshift is a data warehouse for analytical queries, not an ingestion service for raw streaming data.

Amazon Kinesis Data Firehose

A fully managed service for delivering real-time streaming data to destinations such as Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and Splunk.

  • Automatically scales to match data throughput.
  • Requires no administration, provisioning, or ongoing maintenance.
  • Supports data transformation and compression before delivery.

Memory trick: Firehose funnels fast data.

More Data Storage and Management questions