AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy

A data science team has developed a new image classification model using a custom TensorFlow container. They need to deploy this model to a SageMaker real-time endpoint for inference. The model artifact is stored in Amazon S3. Which component is primarily responsible for serving the model predictions at the endpoint?

  1. ASageMaker Inference Container
  2. BSageMaker Processing Job
  3. CSageMaker Feature Store
  4. DSageMaker Ground Truth
Show answer & explanation

Correct answer: A. SageMaker Inference Container

The SageMaker Inference Container is responsible for loading the model artifact, running the inference code, and serving predictions when invoked through the real-time endpoint.

Why the other options are wrong

  • B. SageMaker Processing Jobs are for data preprocessing, post-processing, and model evaluation, not for real-time inference serving.
  • C. SageMaker Feature Store is for managing and serving features for training and inference, not for model serving itself.
  • D. SageMaker Ground Truth is a data labeling service, unrelated to model deployment or inference serving.

SageMaker Inference Container

A Docker container orchestrated by Amazon SageMaker to host and serve machine learning models for real-time or batch inference.

  • Contains model serving code (e.g., Flask, Gunicorn)
  • Loads model artifacts from S3
  • Executes prediction logic when invoked
  • Can be custom or built-in SageMaker images

Memory trick: Containers Run Models Efficiently.

More Machine Learning Implementation and Operations questions