AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy
A financial institution is deploying a new fraud detection model. The model is highly accurate but requires a very low-latency response time for real-time transactions. The current deployment strategy uses a single SageMaker endpoint. To ensure high availability and minimize latency spikes during peak loads, the institution wants to implement a strategy that automatically adjusts the number of inference instances based on traffic while distributing requests across multiple Availability Zones. Which SageMaker feature should be used?
- ASageMaker Neo for model compilation and optimization.
- BSageMaker Endpoint Auto Scaling with Multi-AZ deployment.
- CSageMaker Batch Transform for high-throughput inference.
- DSageMaker Model Monitor for anomaly detection.
Show answer & explanationAnswer & explanation
Correct answer: B. SageMaker Endpoint Auto Scaling with Multi-AZ deployment.
SageMaker Endpoint Auto Scaling automatically adjusts the number of instances based on traffic, while Multi-AZ deployment ensures high availability and distributes requests across multiple Availability Zones, meeting the low-latency and high-availability requirements.
Why the other options are wrong
- A. SageMaker Neo optimizes models for specific hardware but does not provide auto-scaling or multi-AZ deployment capabilities for endpoints.
- C. SageMaker Batch Transform is designed for asynchronous, large-volume inference on entire datasets, not for real-time low-latency predictions.
- D. SageMaker Model Monitor is used for detecting data and model quality issues, not for managing endpoint scaling or availability.
SageMaker Endpoint Auto Scaling
SageMaker Endpoint Auto Scaling automatically adjusts the number of inference instances for a SageMaker endpoint based on demand, ensuring optimal performance and cost efficiency.
- Dynamically scales instances up or down.
- Supports both CPU and GPU instances.
- Can be configured with custom scaling policies.
Memory trick: Auto-scaling endpoints keep your ML reliable and ready for anything.