AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy
A large e-commerce company uses an ML model for real-time product recommendations. The model is deployed on a SageMaker endpoint that experiences highly variable traffic patterns throughout the day, with significant spikes during peak shopping seasons and promotional events. To ensure cost-effectiveness and high availability, the solution must automatically scale the endpoint's inference capacity up and down based on demand. Which SageMaker feature should be configured?
- ASageMaker Elastic Inference
- BSageMaker Batch Transform
- CSageMaker Model Monitor
- DSageMaker Endpoint Auto Scaling
Show answer & explanationAnswer & explanation
Correct answer: D. SageMaker Endpoint Auto Scaling
SageMaker Endpoint Auto Scaling allows you to automatically adjust the number of instances provisioned for a SageMaker endpoint based on predefined metrics (like CPU utilization or custom inference metrics) and target values. This ensures that the endpoint can handle variable traffic while optimizing costs.
Why the other options are wrong
- A. SageMaker Elastic Inference adds GPU acceleration to CPU instances, but it does not automatically scale the number of instances based on traffic.
- B. SageMaker Batch Transform is for processing large datasets asynchronously, not for real-time, auto-scaling inference.
- C. SageMaker Model Monitor is for detecting model and data quality issues, not for scaling inference capacity.
SageMaker Endpoint Auto Scaling
A feature that automatically adjusts the number of instances for a SageMaker real-time endpoint based on predefined scaling policies and metrics.
- Optimizes cost by scaling down during low demand.
- Ensures high availability and performance during peak demand.
- Uses CloudWatch metrics to trigger scaling actions.
Memory trick: Auto Scaling is the endpoint's flexible workforce.