AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard

A machine learning engineer is tasked with optimizing the inference performance of a large, complex model deployed on a SageMaker endpoint. The model has many layers, and the bottleneck is identified as the sequential execution of these layers. The engineer wants to parallelize parts of the inference process to reduce latency. Which advanced inference optimization technique is most appropriate for this scenario?

  1. AQuantization
  2. BBatch inference with larger batch sizes
  3. CModel parallelism
  4. DData parallelism
Show answer & explanation

Correct answer: C. Model parallelism

Model parallelism is used when a single model is too large to fit into memory or to be processed efficiently by a single device. It involves splitting the model's layers or operations across multiple devices (e.g., GPUs) and processing them in parallel, which is suitable for reducing latency in complex, sequential models.

Why the other options are wrong

  • A. Quantization reduces model size and memory footprint but doesn't inherently parallelize sequential layer execution to reduce latency in the described manner.
  • B. Batch inference increases throughput but also increases latency for individual requests, as requests wait for a full batch.
  • D. Data parallelism splits the *data* across multiple devices, with each device processing a full copy of the model, which is more for throughput than reducing single-request latency of a complex model.

Model Parallelism (Inference)

An inference optimization technique where a single, large model is partitioned across multiple compute devices (e.g., GPUs), with different parts of the model (e.g., layers) processed in parallel.

  • Reduces latency for very large and complex models.
  • Allows models that exceed single-device memory to be deployed.
  • Contrast with data parallelism, which distributes data, not the model.

Memory trick: Model parallelism splits the brain, data parallelism splits the workload.

More Machine Learning Implementation and Operations questions