AWS Certified Machine Learning – SpecialtyData EngineeringHard
A startup is building a recommendation engine and needs to process petabytes of clickstream data from its website. This data arrives continuously, and transformations, such as sessionization and feature extraction, must be applied in near real-time before being stored for model training. The solution must be serverless, highly scalable, and cost-effective. Which AWS service combination should the startup use to achieve this?
- AAmazon MSK, Amazon EC2, and Amazon EFS
- BAmazon Kinesis Data Streams, AWS Lambda, and Amazon S3
- CAmazon Kinesis Firehose, AWS Glue, and Amazon RDS
- DAmazon SQS, AWS Batch, and Amazon Redshift
Show answer & explanationAnswer & explanation
Correct answer: B. Amazon Kinesis Data Streams, AWS Lambda, and Amazon S3
Amazon Kinesis Data Streams is ideal for capturing and processing large streams of data in real-time. AWS Lambda can be used to process records from Kinesis Data Streams, applying transformations like sessionization and feature extraction in a serverless and scalable manner. The processed data can then be stored in Amazon S3, which serves as a cost-effective and highly scalable data lake for subsequent model training.
Why the other options are wrong
- A. Amazon MSK (Managed Streaming for Apache Kafka) requires managing Kafka clusters. Amazon EC2 instances require provisioning and management. Amazon EFS is a file system, not ideal for petabyte-scale object storage for ML.
- C. Amazon Kinesis Firehose delivers streams to destinations but has limited real-time processing capabilities compared to Data Streams with Lambda. AWS Glue is for ETL, typically batch-oriented, not near real-time. Amazon RDS is not suitable for petabyte-scale raw data storage.
- D. Amazon SQS is a message queuing service, not designed for streaming petabytes of data. AWS Batch is for batch processing, not real-time. Amazon Redshift is a data warehouse, not a primary storage for raw, continuously arriving data.
Real-time Serverless Data Ingestion & Processing
A common AWS architectural pattern for handling high-volume, continuous data streams, transforming them in near real-time, and storing them for machine learning workloads.
- Uses Kinesis Data Streams for data ingestion.
- Leverages AWS Lambda for serverless, event-driven processing.
- Stores processed data in Amazon S3 for data lake capabilities.
- Scalable, cost-effective, and fully managed components.
Memory trick: Stream with Kinesis, Lambda processes, S3 stores the ML insights.