A global media streaming service is building a new recommendation engine that analyzes user viewing habits and content metadata to suggest personalized content. The engine needs to process petabytes of historical data for model training and then serve real-time recommendations with low latency (under 100ms) to millions of concurrent users. The data processing for model training can be batch-oriented, but the inference must be highly responsive. Which architecture should the company choose for this recommendation engine?
- AAmazon Redshift for historical data analysis and Amazon ElastiCache for Redis for real-time recommendations.
- BAmazon EMR for batch processing of historical data, Amazon Kinesis Data Streams for real-time data ingestion, and AWS Lambda for model inference.
- CAmazon S3 for raw data storage, AWS Glue for ETL and feature engineering, Amazon SageMaker for model training and hosting, and Amazon DynamoDB for storing real-time inference results.
- DAmazon OpenSearch Service for historical data indexing and analysis, and Amazon Neptune for real-time graph-based recommendations.
Show answer & explanationAnswer & explanation
Correct answer: C. Amazon S3 for raw data storage, AWS Glue for ETL and feature engineering, Amazon SageMaker for model training and hosting, and Amazon DynamoDB for storing real-time inference results.
This architecture leverages S3 for scalable and cost-effective raw data storage. AWS Glue is excellent for serverless ETL and feature engineering on petabyte-scale data for model training. Amazon SageMaker provides a comprehensive platform for building, training (batch-oriented), and deploying (real-time endpoint hosting) machine learning models. DynamoDB is a highly scalable, low-latency NoSQL database suitable for storing and serving real-time inference results for millions of concurrent users with single-digit millisecond latency.
Why the other options are wrong
- A. Redshift is good for historical data, but ElastiCache for Redis is primarily a caching solution, not a persistent store for real-time recommendations, nor does it provide ML model hosting.
- B. EMR is suitable for batch processing. Kinesis can ingest real-time data, but using AWS Lambda for direct model inference might introduce latency issues for complex models and isn't as optimized for ML inference as SageMaker endpoints.
- D. OpenSearch Service is for search and analytics, not typically the primary store for petabyte-scale raw data for ML training. Neptune is a graph database, useful for certain recommendation types, but doesn't cover the full scope of batch training and general real-time serving as comprehensively as SageMaker and DynamoDB.
ML Recommendation Engine Architecture
A system designed to generate personalized content suggestions, often involving separate components for batch model training on historical data and low-latency real-time inference.
- Separates batch training from real-time inference.
- Uses scalable storage for raw data (S3).
- Leverages SageMaker for comprehensive ML lifecycle management.
Memory trick: S3 stores the past, Glue transforms, SageMaker learns, DynamoDB serves the future.