AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium
A large e-commerce platform uses an ML model for real-time product recommendations. The model processes millions of inference requests per hour, with traffic patterns that fluctuate significantly throughout the day. To ensure responsiveness and optimize costs, the platform needs to automatically scale the SageMaker endpoint instances based on the incoming traffic load. Which SageMaker feature should be configured?
- ASageMaker Endpoint Auto Scaling
- BSageMaker Model Monitor
- CSageMaker Batch Transform
- DSageMaker Elastic Inference
Show answer & explanationAnswer & explanation
Correct answer: A. SageMaker Endpoint Auto Scaling
SageMaker Endpoint Auto Scaling automatically adjusts the number of instances for a SageMaker endpoint based on predefined metrics (like CPU utilization or invocation count) and target values. This ensures that the endpoint can handle fluctuating traffic loads efficiently, maintaining responsiveness while optimizing costs by scaling instances up or down as needed.
Why the other options are wrong
- B. SageMaker Model Monitor tracks model quality and data drift, not endpoint instance scaling.
- C. SageMaker Batch Transform is for offline, asynchronous inference and does not manage real-time endpoint scaling.
- D. SageMaker Elastic Inference attaches GPU acceleration to CPU-based instances, optimizing cost/performance for specific model types, but does not auto-scale the number of instances itself.
SageMaker Endpoint Auto Scaling
A SageMaker feature that automatically adjusts the number of inference instances for a real-time endpoint based on predefined scaling policies and metrics.
- Scales up to handle increased traffic.
- Scales down to reduce costs during low traffic.
- Uses CloudWatch metrics (e.g., CPU, invocations) to trigger scaling events.
Memory trick: Auto-Scale Smart, Cost and Speed at Heart.