AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A gaming company collects telemetry data from millions of players globally. This data includes player actions, in-game events, and device information, arriving at extremely high velocity and volume. The data needs to be ingested, transformed, and made available for real-time analytics and machine learning model training to detect anomalies and personalize player experiences. The solution must handle petabytes of data, offer high availability, and be able to scale dynamically. Which AWS service is best suited for the initial ingestion and buffering of this high-velocity streaming data?

  1. AAWS DataSync
  2. BAmazon SQS
  3. CAmazon Kinesis Data Firehose
  4. DAmazon Redshift
Show answer & explanation

Correct answer: C. Amazon Kinesis Data Firehose

Amazon Kinesis Data Firehose is a fully managed service that reliably loads streaming data into data lakes, data stores, and analytics services. It automatically scales to match the throughput of your data and can apply basic transformations (e.g., format conversion, compression) before delivery, making it ideal for high-velocity ingestion and buffering.

Why the other options are wrong

  • A. AWS DataSync is a data transfer service that simplifies, automates, and accelerates moving data between on-premises storage and AWS storage services, or between AWS storage services. It's for bulk data movement, not real-time streaming ingestion.
  • B. Amazon SQS (Simple Queue Service) is a message queuing service, suitable for decoupling microservices or managing task queues, but not designed for high-throughput, continuous streaming data ingestion for analytics and ML.
  • D. Amazon Redshift is a petabyte-scale data warehouse for analytical queries. It's a destination for processed data, not an ingestion service for raw, high-velocity streaming data.

Amazon Kinesis Data Firehose

A fully managed service for delivering real-time streaming data to destinations like Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and Splunk, with built-in buffering and transformation.

  • Fully managed and automatically scales.
  • No servers to provision or manage.
  • Buffering and batching capabilities.
  • Supports basic transformations (e.g., format conversion, compression).
  • Cost-effective for loading data into data lakes/warehouses.

Memory trick: Firehose funnels data streams to storage, fast and easy.

More Data Engineering questions