AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium
A large e-commerce company uses an ML model for real-time fraud detection. The model processes millions of transactions per hour, and high availability is paramount. Any downtime or performance degradation could lead to significant financial losses. The company needs to ensure that the inference endpoint remains operational and performs optimally even during AWS Availability Zone outages or instance failures. Which SageMaker deployment configuration best addresses this requirement?
- ADeploy the model using SageMaker Batch Transform for high throughput and fault tolerance.
- BDeploy the model to a single SageMaker endpoint with multiple instances in a single Availability Zone.
- CDeploy the model to a SageMaker endpoint configured with multiple production variants, each in a different AWS Region.
- DDeploy the model to a SageMaker endpoint configured with multiple instances distributed across multiple Availability Zones.
Show answer & explanationAnswer & explanation
Correct answer: D. Deploy the model to a SageMaker endpoint configured with multiple instances distributed across multiple Availability Zones.
Deploying a SageMaker endpoint with multiple instances distributed across multiple Availability Zones is the standard AWS practice for achieving high availability and fault tolerance. If one AZ experiences an outage or an instance fails, traffic is automatically routed to healthy instances in other AZs, ensuring continuous operation and minimal impact on performance.
Why the other options are wrong
- A. SageMaker Batch Transform is for offline, non-real-time inference and does not meet the requirement for real-time fraud detection with high availability.
- B. Deploying to a single AZ, even with multiple instances, does not protect against an AZ-wide outage.
- C. Deploying to multiple AWS Regions provides disaster recovery but is overkill for AZ outages and introduces higher latency and complexity than multi-AZ within a single region.
SageMaker Multi-AZ Endpoint
A SageMaker inference endpoint configured with multiple instances deployed across different AWS Availability Zones for high availability and fault tolerance.
- Protects against AZ outages and instance failures
- Ensures continuous operation and minimal downtime
- Traffic is automatically routed to healthy instances
- Achieved by specifying multiple instances for an endpoint config
Memory trick: Distribute for resilience, keep the ML flowing.