AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard

A pharmaceutical company is developing a machine learning model to assist in drug discovery. The model is computationally intensive, requiring significant GPU resources for inference, but the inference requests are infrequent and latency is not extremely critical (a few seconds is acceptable). The company wants to optimize costs for these inference workloads. Which SageMaker feature should they consider to achieve cost-effective GPU-powered inference?

  1. ASageMaker Multi-Model Endpoints
  2. BSageMaker Elastic Inference
  3. CSageMaker Serverless Inference
  4. DSageMaker Batch Transform
Show answer & explanation

Correct answer: B. SageMaker Elastic Inference

SageMaker Elastic Inference allows you to attach GPU acceleration to CPU-based EC2 instances or SageMaker instances at a fraction of the cost of a full GPU instance. Since the model requires GPU resources but inference is infrequent and latency is not extremely critical, Elastic Inference provides a cost-effective way to get GPU acceleration without paying for oversized GPU instances, aligning perfectly with the cost optimization goal for infrequent workloads.

Why the other options are wrong

  • A. Multi-Model Endpoints allow hosting multiple models on a single endpoint, reducing operational overhead, but don't specifically address cost-effective GPU acceleration.
  • C. Serverless Inference is for models with sparse traffic and bursty workloads, scaling to zero, but it does not support GPU acceleration directly for custom models.
  • D. Batch Transform is for offline, batch processing of large datasets, not suitable for real-time (even if infrequent) inference where a few seconds latency is acceptable.

SageMaker Elastic Inference

A feature that allows attaching GPU-powered inference acceleration to Amazon EC2 and SageMaker instances, reducing the cost of deep learning inference.

  • Provides GPU acceleration at a lower cost than full GPU instances.
  • Suitable for models requiring GPU but don't fully utilize a dedicated GPU.
  • Integrates with existing CPU instances.

Memory trick: Elastic inference, flexible spend.

More Machine Learning Implementation and Operations questions