AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard
A machine learning engineer is deploying a large language model (LLM) on Amazon SageMaker. The model is too large to fit into the memory of a single GPU instance for inference. The engineer needs to optimize inference performance while ensuring the model can be served efficiently. Which SageMaker feature should the engineer use to distribute the model across multiple GPUs or instances for inference?
- ASageMaker Neo
- BSageMaker Inference Recommender
- CSageMaker Model Parallelism (for inference)
- DSageMaker Batch Transform
Show answer & explanationAnswer & explanation
Correct answer: C. SageMaker Model Parallelism (for inference)
SageMaker Model Parallelism for inference allows large models, such as LLMs, to be split and distributed across multiple GPUs within a single instance or across multiple instances. This enables serving models that exceed the memory capacity of a single device and optimizes inference performance through parallel execution.
Why the other options are wrong
- A. SageMaker Neo optimizes models for specific hardware targets but doesn't inherently distribute a single model across multiple devices for inference.
- B. SageMaker Inference Recommender helps choose the best instance type and configuration for a model but doesn't directly provide model parallelism for large models.
- D. SageMaker Batch Transform is for processing large datasets asynchronously, not for real-time, distributed inference of a single large model.
SageMaker Model Parallelism (Inference)
A technique to distribute a single large machine learning model's layers or components across multiple GPUs or instances to overcome memory constraints and accelerate inference.
- Crucial for deploying very large models (e.g., LLMs)
- Splits model layers or parameters across devices
- Reduces memory footprint per device
- Can improve throughput and reduce latency for large models
Memory trick: Parallel Models Split Big Brains.