AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium
A pharmaceutical company is developing a machine learning model to assist in drug discovery. The model is a large transformer-based architecture deployed on an Amazon SageMaker endpoint. While the model provides high accuracy, the inference latency is a critical concern, and the GPU instances required are expensive. The team wants to optimize the inference cost and latency without compromising model accuracy. Which SageMaker feature should they use to achieve this?
- ASageMaker Elastic Inference
- BSageMaker Neo
- CSageMaker Model Monitor
- DSageMaker Feature Store
Show answer & explanationAnswer & explanation
Correct answer: A. SageMaker Elastic Inference
SageMaker Elastic Inference allows you to attach GPU-powered inference acceleration to Amazon EC2 instances or SageMaker instances. This enables you to choose a CPU instance type that meets your memory requirements and then attach the right amount of GPU acceleration (Elastic Inference Accelerator) to improve inference performance and reduce costs, as you only pay for the fraction of GPU power you need, rather than a full GPU instance.
Why the other options are wrong
- B. SageMaker Neo compiles models to optimize them for specific hardware targets, which can improve performance, but Elastic Inference specifically addresses attaching GPU acceleration to CPU instances for cost-effective GPU inference.
- C. SageMaker Model Monitor is for monitoring model quality and drift in production, not for optimizing inference cost or latency.
- D. SageMaker Feature Store is for storing and serving features for training and inference, not for optimizing model inference performance or cost.
SageMaker Elastic Inference
An Amazon SageMaker feature that allows you to attach GPU-powered inference acceleration to Amazon EC2 and SageMaker instances, enabling cost-effective deep learning inference.
- Decouples GPU acceleration from instance type.
- Reduces inference costs by paying only for required GPU acceleration.
- Improves latency for deep learning models.
- Supports various deep learning frameworks.
Memory trick: Elastic Inference: Stretch GPU power to fit, save cash.