AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium
An e-commerce company uses an ML model for real-time product recommendations. The model processes millions of inference requests per hour, with traffic fluctuating significantly throughout the day. To manage costs and ensure consistent performance, the ML team needs to dynamically adjust the number of inference instances based on the incoming request load. Which Amazon SageMaker feature should they implement?
- ASageMaker Elastic Inference
- BSageMaker Model Monitor
- CSageMaker Pipelines
- DSageMaker Auto Scaling
Show answer & explanationAnswer & explanation
Correct answer: D. SageMaker Auto Scaling
SageMaker Auto Scaling automatically adjusts the number of instances provisioned for a SageMaker endpoint based on predefined scaling policies and metrics (like CPU utilization or invocation per minute). This allows the e-commerce company to efficiently manage costs by scaling down during low traffic and maintain performance by scaling up during peak loads.
Why the other options are wrong
- A. SageMaker Elastic Inference is for attaching GPU acceleration to EC2 and SageMaker instances to reduce inference costs for deep learning models, not for scaling the number of instances.
- B. SageMaker Model Monitor is for detecting data and model quality drift, not for dynamic instance scaling.
- C. SageMaker Pipelines are for orchestrating ML workflows, not for dynamic scaling of deployed models.
SageMaker Auto Scaling
A feature of Amazon SageMaker that automatically adjusts the number of instances provisioned for a SageMaker endpoint in response to changes in workload.
- Optimizes costs by scaling instances up and down.
- Maintains performance during fluctuating traffic.
- Uses CloudWatch metrics and scaling policies.
Memory trick: Auto-scale for happy wallets and users.