AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium

An e-commerce company uses a fraud detection model that needs to process millions of transactions daily. The model's predictions are critical, and any downtime or high latency could result in significant financial losses. The company requires a highly available and fault-tolerant deployment strategy for its SageMaker endpoint, ensuring continuous operation even during instance failures or updates. Which SageMaker endpoint configuration best meets these requirements?

  1. ASingle-instance endpoint with a basic auto-scaling policy.
  2. BBatch transform job on a scheduled basis.
  3. CServerless inference endpoint with provisioned concurrency.
  4. DMulti-instance endpoint across multiple Availability Zones with auto-scaling.
Show answer & explanation

Correct answer: D. Multi-instance endpoint across multiple Availability Zones with auto-scaling.

Deploying a multi-instance SageMaker endpoint across multiple Availability Zones (AZs) with auto-scaling provides high availability and fault tolerance. If an instance or an entire AZ fails, traffic is routed to healthy instances in other AZs, ensuring continuous operation.

Why the other options are wrong

  • A. A single-instance endpoint is a single point of failure and does not provide high availability.
  • B. Batch transform is for offline processing and cannot be used for real-time, low-latency requirements.
  • C. Serverless inference offers scalability and cost efficiency for intermittent traffic but might introduce cold starts which could impact low-latency requirements for critical applications, unless provisioned concurrency is carefully managed and scaled.

SageMaker Multi-AZ Endpoint

A SageMaker real-time endpoint configuration that deploys instances across multiple AWS Availability Zones (AZs) for high availability and fault tolerance.

  • Protects against single point of failure in an AZ.
  • Combines with auto-scaling for dynamic capacity management.
  • Essential for critical, low-latency production workloads.

Memory trick: Multiple instances in multiple zones keep the model alive.

More Machine Learning Implementation and Operations questions