AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A global e-commerce company needs to track user behavior on its website and mobile applications. This generates billions of small (approx. 500-byte) JSON events daily, with unpredictable spikes in traffic. The data needs to be delivered to an Amazon S3 data lake for batch analytics and also to an Amazon Redshift data warehouse for reporting dashboards. The company wants a fully managed solution with minimal setup and operational overhead, and it needs basic data transformation (e.g., converting JSON to Parquet) before delivery. Which AWS service is most appropriate for this ingestion and delivery?

  1. AAWS Glue Streaming ETL
  2. BAmazon Kinesis Data Firehose
  3. CAmazon Kinesis Data Streams
  4. DAWS Transfer Family
Show answer & explanation

Correct answer: B. Amazon Kinesis Data Firehose

Amazon Kinesis Data Firehose is a fully managed service for delivering real-time streaming data to destinations like Amazon S3 and Amazon Redshift. It automatically scales to handle traffic spikes, requires minimal setup, and can perform basic transformations like JSON to Parquet conversion before delivery, making it ideal for this use case with minimal operational overhead.

Why the other options are wrong

  • A. AWS Glue Streaming ETL is for complex, continuous real-time transformations, which is overkill for 'basic data transformation' and adds more complexity than Firehose for simple delivery to S3 and Redshift.
  • C. Kinesis Data Streams is a low-level streaming service that requires consumers to be built and managed, adding operational overhead not desired for 'minimal setup and operational overhead' and 'basic data transformation'.
  • D. AWS Transfer Family is for transferring files over SFTP/FTPS/FTP, not for ingesting high-velocity, real-time event streams from web/mobile applications.

Amazon Kinesis Data Firehose

A fully managed service for delivering real-time streaming data to destinations such as Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and Splunk.

  • Fully managed and scales automatically.
  • Supports various destinations (S3, Redshift, etc.).
  • Offers built-in data transformation (e.g., format conversion, Lambda for custom logic).
  • Batching, compression, and encryption are handled automatically.

Memory trick: Firehose 'sprays' your data to multiple destinations.

More Data Ingestion and Transformation questions