AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium
A pharmaceutical company is developing a machine learning model to assist in drug discovery. The model is computationally intensive and requires significant GPU resources for inference. To optimize costs while maintaining an acceptable latency of less than 100ms, they want to use a shared, elastic GPU resource for inference that can be attached to their CPU-based SageMaker instances. Which SageMaker feature should they use?
- ASageMaker Multi-Model Endpoints
- BSageMaker Elastic Inference
- CSageMaker Asynchronous Inference
- DSageMaker Inference Recommender
Show answer & explanationAnswer & explanation
Correct answer: B. SageMaker Elastic Inference
SageMaker Elastic Inference allows attaching GPU acceleration to CPU-based SageMaker instances, providing cost-effective GPU inference. This is ideal for models that require GPU acceleration but don't need a full GPU instance, helping to optimize costs while meeting latency requirements for computationally intensive models.
Why the other options are wrong
- A. SageMaker Multi-Model Endpoints allow deploying multiple models on a single endpoint, reducing costs by sharing resources, but it doesn't provide elastic GPU acceleration to CPU instances.
- C. SageMaker Asynchronous Inference is for large payloads or long-running inferences where real-time latency is not critical, which contradicts the 100ms latency requirement.
- D. SageMaker Inference Recommender helps choose the best instance type and configuration for deployment, but it doesn't provide elastic GPU acceleration for CPU instances.
SageMaker Elastic Inference
A SageMaker feature that allows attaching GPU acceleration to CPU-based SageMaker instances, providing cost-effective inference for deep learning models.
- Reduces inference costs by up to 75%
- Supports various deep learning frameworks
- Provides flexible GPU acceleration without full GPU instances
- Ideal for models needing some GPU power but not a dedicated GPU instance
Memory trick: Elastic Inference: GPU power without the full price tag.