AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard

A machine learning engineer is deploying a large language model (LLM) on Amazon SageMaker for real-time inference. The LLM is several tens of billions of parameters, making it too large to fit into the memory of a single GPU instance or even multiple GPUs on a single instance without significant performance degradation. The engineer needs to optimize the inference performance and reduce latency by distributing the model across multiple GPU instances. Which SageMaker optimization technique is designed for this scenario?

  1. ASageMaker Multi-Model Endpoints
  2. BSageMaker Model Parallelism (Inference)
  3. CSageMaker Elastic Inference
  4. DSageMaker Neo
Show answer & explanation

Correct answer: B. SageMaker Model Parallelism (Inference)

SageMaker Model Parallelism (Inference) allows large models (like LLMs) that exceed the memory capacity of a single GPU or instance to be split across multiple GPUs or instances. This technique distributes the model's layers and parameters, enabling inference for very large models that would otherwise be infeasible or suffer from high latency on smaller hardware.

Why the other options are wrong

  • A. SageMaker Multi-Model Endpoints host multiple models on a single endpoint, reducing cost for many small models, not for splitting one very large model across multiple instances.
  • C. SageMaker Elastic Inference adds GPU acceleration to CPU instances for smaller models, but it's not designed for models that *must* be split across multiple instances due to their sheer size.
  • D. SageMaker Neo optimizes models for specific hardware architectures (e.g., edge devices) by compiling them, but it does not address the challenge of splitting a model across multiple instances for inference.

SageMaker Model Parallelism (Inference)

An advanced inference optimization technique that splits a large machine learning model (e.g., LLM) across multiple GPUs or instances to overcome memory limitations and improve inference latency.

  • Crucial for deploying models too large for single-device memory.
  • Distributes model layers/parameters across hardware.
  • Reduces inference latency for very large models.

Memory trick: Parallelism makes giant models fly fast.

More Machine Learning Implementation and Operations questions