AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy

A large e-commerce platform uses a recommendation engine powered by a machine learning model deployed on Amazon SageMaker. During peak shopping seasons, the inference requests can surge unexpectedly, leading to increased latency and potential service disruptions. The platform needs a solution that can automatically adjust the endpoint's capacity to handle these unpredictable traffic spikes without manual intervention and ensure consistent low latency. Which combination of SageMaker features should be implemented?

  1. ASageMaker Model Monitor and SageMaker Neo.
  2. BSageMaker Endpoint Auto Scaling and Multi-AZ deployment.
  3. CSageMaker Data Wrangler and SageMaker Pipelines.
  4. DSageMaker Batch Transform and Elastic Inference.
Show answer & explanation

Correct answer: B. SageMaker Endpoint Auto Scaling and Multi-AZ deployment.

SageMaker Endpoint Auto Scaling dynamically adjusts the number of inference instances based on load, ensuring capacity during traffic spikes. Multi-AZ deployment distributes these instances across multiple Availability Zones, providing high availability and resilience against single-point failures, which is crucial for consistent low latency and service continuity.

Why the other options are wrong

  • A. Model Monitor detects data/model quality issues, and Neo optimizes models for specific hardware; neither directly addresses dynamic scaling for unpredictable traffic spikes.
  • C. Data Wrangler prepares data, and SageMaker Pipelines orchestrates ML workflows; neither directly manages real-time endpoint capacity or availability.
  • D. Batch Transform is for offline, large-scale inference, not real-time, and Elastic Inference is for attaching GPU acceleration to CPU instances, not for dynamic capacity scaling.

SageMaker Auto Scaling & Multi-AZ

SageMaker Endpoint Auto Scaling dynamically adjusts inference instance counts based on demand, while Multi-AZ deployment distributes instances across Availability Zones for high availability and fault tolerance.

  • Auto Scaling adjusts capacity for traffic changes.
  • Multi-AZ ensures high availability and resilience.
  • Together, they provide robust, scalable, and fault-tolerant inference.

Memory trick: Scaling up and staying up is key for peak ML performance.

More Machine Learning Implementation and Operations questions