AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium
A large e-commerce platform uses an ML model for real-time fraud detection. The model processes millions of transactions per hour, and high availability is paramount. Any downtime or performance degradation could lead to significant financial losses. The team needs to deploy the model in a way that provides automatic failover and resilience against Availability Zone (AZ) outages. Which SageMaker deployment option should they choose?
- ADeploy to a SageMaker Multi-AZ endpoint.
- BDeploy multiple single-AZ endpoints and manage failover with Route 53.
- CDeploy to a SageMaker serverless endpoint.
- DDeploy to a single SageMaker endpoint with auto-scaling enabled.
Show answer & explanationAnswer & explanation
Correct answer: A. Deploy to a SageMaker Multi-AZ endpoint.
A SageMaker Multi-AZ endpoint automatically distributes inference requests across instances in multiple Availability Zones. If an AZ becomes unhealthy, traffic is automatically routed to healthy instances in other AZs, providing built-in high availability and resilience against AZ outages without manual intervention or additional services like Route 53 for failover management.
Why the other options are wrong
- B. While possible, manually configuring multiple single-AZ endpoints and managing failover with Route 53 is more complex and less integrated than SageMaker's native Multi-AZ endpoint feature.
- C. SageMaker serverless endpoints are suitable for infrequent, non-latency-critical workloads, not for high-volume, low-latency, and high-availability fraud detection systems.
- D. Auto-scaling on a single-AZ endpoint helps with traffic spikes but does not protect against an entire AZ outage.
SageMaker Multi-AZ Endpoint
A SageMaker deployment option that automatically distributes model instances across multiple AWS Availability Zones, providing high availability and automatic failover for real-time inference.
- Ensures resilience against Availability Zone outages.
- Automatically routes traffic to healthy instances.
- Provides built-in high availability for critical models.
- Simplifies deployment compared to manual multi-AZ setups.
Memory trick: Multi-AZ: Spread your models across regions, never fail.