AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy

A data science team is developing a critical machine learning model for real-time anomaly detection. They need to ensure that the model is always available and highly resilient to instance failures or Availability Zone (AZ) outages. The application requires very low latency for predictions. Which SageMaker endpoint configuration best addresses these requirements?

  1. ADeploying the model to a multi-instance SageMaker endpoint within a single Availability Zone.
  2. BDeploying the model to a single-instance SageMaker endpoint in a single Availability Zone.
  3. CDeploying the model to a SageMaker serverless endpoint for automatic scaling.
  4. DDeploying the model to a multi-instance SageMaker endpoint across multiple Availability Zones.
Show answer & explanation

Correct answer: D. Deploying the model to a multi-instance SageMaker endpoint across multiple Availability Zones.

Deploying a multi-instance SageMaker endpoint across multiple Availability Zones provides the highest level of availability and resilience. If one instance or an entire AZ fails, traffic is automatically routed to healthy instances in other AZs, ensuring continuous operation and low latency.

Why the other options are wrong

  • A. Multiple instances in a single AZ provide some resilience to instance failure but not to an AZ outage.
  • B. A single instance in a single AZ is not highly available or resilient to failures.
  • C. SageMaker Serverless Endpoints offer automatic scaling and cost optimization for infrequent traffic but might introduce cold start latency and are not explicitly designed for the highest resilience against AZ outages like multi-AZ deployments.

SageMaker Multi-AZ Endpoint

A SageMaker real-time endpoint configured to distribute inference instances across multiple AWS Availability Zones.

  • Provides high availability and fault tolerance against AZ outages.
  • Ensures continuous operation for critical real-time ML applications.
  • Traffic is automatically routed to healthy instances.

Memory trick: Spread your instances far and wide, so nothing can hide!

More Machine Learning Implementation and Operations questions