AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy
A large e-commerce platform uses a recommendation engine powered by a machine learning model deployed on Amazon SageMaker. During peak shopping seasons, the inference requests can surge unexpectedly, leading to increased latency and potential service disruptions. The platform needs a solution that can automatically adjust the endpoint's capacity to handle these unpredictable traffic spikes without manual intervention and ensure consistent low latency. Which combination of SageMaker features should be implemented?
- ASageMaker Model Monitor and SageMaker Neo.
- BSageMaker Endpoint Auto Scaling and Multi-AZ deployment.
- CSageMaker Data Wrangler and SageMaker Pipelines.
- DSageMaker Batch Transform and Elastic Inference.
Show answer & explanationAnswer & explanation
Correct answer: B. SageMaker Endpoint Auto Scaling and Multi-AZ deployment.
SageMaker Endpoint Auto Scaling dynamically adjusts the number of inference instances based on load, ensuring capacity during traffic spikes. Multi-AZ deployment distributes these instances across multiple Availability Zones, providing high availability and resilience against single-point failures, which is crucial for consistent low latency and service continuity.
Why the other options are wrong
- A. Model Monitor detects data/model quality issues, and Neo optimizes models for specific hardware; neither directly addresses dynamic scaling for unpredictable traffic spikes.
- C. Data Wrangler prepares data, and SageMaker Pipelines orchestrates ML workflows; neither directly manages real-time endpoint capacity or availability.
- D. Batch Transform is for offline, large-scale inference, not real-time, and Elastic Inference is for attaching GPU acceleration to CPU instances, not for dynamic capacity scaling.
SageMaker Auto Scaling & Multi-AZ
SageMaker Endpoint Auto Scaling dynamically adjusts inference instance counts based on demand, while Multi-AZ deployment distributes instances across Availability Zones for high availability and fault tolerance.
- Auto Scaling adjusts capacity for traffic changes.
- Multi-AZ ensures high availability and resilience.
- Together, they provide robust, scalable, and fault-tolerant inference.
Memory trick: Scaling up and staying up is key for peak ML performance.