AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard

A pharmaceutical company is developing a machine learning model to assist in drug discovery, which involves extremely large and complex deep learning models. The inference latency is critical, but the cost of deploying full GPU instances for every inference request is prohibitive. The model needs to run on SageMaker, and the team wants to optimize cost-performance for inference by leveraging GPU acceleration only when necessary, attached to CPU instances. Which SageMaker feature should they utilize?

  1. ASageMaker Batch Transform
  2. BSageMaker Inference Recommender
  3. CSageMaker Elastic Inference
  4. DSageMaker Multi-AZ Endpoint
Show answer & explanation

Correct answer: C. SageMaker Elastic Inference

SageMaker Elastic Inference (EI) allows you to attach fractional GPU acceleration to CPU-based SageMaker instances or EC2 instances. This is ideal for deep learning models where the GPU is not fully utilized during inference, as it provides a cost-effective way to get GPU acceleration without provisioning full GPU instances. This directly addresses the need for critical inference latency with cost optimization for models that don't need dedicated, full GPU power.

Why the other options are wrong

  • A. Batch Transform is for offline inference and does not address real-time latency or cost-optimized GPU acceleration.
  • B. Inference Recommender helps choose the best instance type for deployment, but it doesn't provide the fractional GPU acceleration mechanism itself.
  • D. Multi-AZ endpoint provides high availability, not cost-optimized GPU acceleration.

SageMaker Elastic Inference

A SageMaker feature that allows attaching fractional GPU acceleration to Amazon EC2 and SageMaker instances, enabling cost-effective deep learning inference.

  • Provides GPU acceleration at a lower cost than full GPU instances.
  • Ideal for models that don't fully utilize a dedicated GPU.
  • Attaches to CPU instances for inference workloads.

Memory trick: Elastic Inference: GPU Power, Feather Light, Cost's Not a Fight.

More Machine Learning Implementation and Operations questions