A global media streaming service is building a new recommendation engine that analyzes user viewing patterns, historical data, and real-time interactions to provide personalized content suggestions. The architecture must handle petabytes of historical data, process real-time clickstream data for immediate recommendations, and support complex machine learning model training and inference at scale. The solution needs to be cost-effective and highly performant. Which AWS architecture provides the best combination of services for this recommendation engine?
- AAmazon S3 for historical data, Amazon Kinesis Data Streams for real-time data, AWS Glue for ETL, Amazon Redshift for model training data, and Amazon SageMaker for model hosting.
- BAmazon S3 for raw historical data, Amazon Kinesis Data Streams for real-time events, Amazon Redshift for data warehousing, Amazon EC2 instances with GPUs for model training, and AWS Lambda for real-time inference.
- CAmazon S3 for raw historical data, AWS Lake Formation for data lake governance, Amazon Kinesis Data Streams for real-time events, Amazon Athena for ad-hoc analysis, Amazon EMR for Spark-based ML training, and Amazon SageMaker for real-time inference.
- DAmazon S3 for historical data, Amazon Kinesis Data Firehose for real-time data ingestion, Amazon EMR for batch processing, Amazon Neptune for graph-based recommendations, and Amazon SageMaker for model deployment.
Show answer & explanationAnswer & explanation
Correct answer: C. Amazon S3 for raw historical data, AWS Lake Formation for data lake governance, Amazon Kinesis Data Streams for real-time events, Amazon Athena for ad-hoc analysis, Amazon EMR for Spark-based ML training, and Amazon SageMaker for real-time inference.
This architecture establishes a robust data lake on S3 with Lake Formation for governance. Kinesis Data Streams handles real-time events. EMR with Spark is ideal for large-scale ML training on petabytes of data, and SageMaker provides scalable, high-performance real-time inference, directly addressing all requirements for a recommendation engine.
Why the other options are wrong
- A. Redshift is a data warehouse, not typically used for direct ML model training on petabytes of raw data. Glue is good for ETL but EMR is more flexible for large-scale Spark ML.
- B. Redshift is not ideal for raw ML training; EC2 instances with GPUs can be used but SageMaker provides a more managed and scalable ML platform. Lambda is not suitable for high-throughput, low-latency real-time inference at scale for complex ML models.
- D. Neptune is for graph databases and might be useful for certain recommendation types, but it's not the primary ML training platform for petabytes of diverse data, and Firehose is less flexible than Kinesis Data Streams for real-time processing.
ML Recommendation Engine Architecture
An AWS architecture for building personalized content recommendation engines, combining data lake capabilities, real-time stream processing, and scalable machine learning platforms.
- Uses S3 for historical data lake.
- Kinesis Data Streams for real-time event ingestion.
- EMR for large-scale ML model training.
- SageMaker for model deployment and inference.
Memory trick: S3, Kinesis, EMR, SageMaker: Recommend the Right Way.