AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy
A data science team has developed a new image classification model using a custom TensorFlow container. They need to deploy this model to a SageMaker real-time endpoint for inference. The model artifact is stored in Amazon S3. Which component is primarily responsible for serving the model predictions at the endpoint?
- ASageMaker Inference Container
- BSageMaker Processing Job
- CSageMaker Feature Store
- DSageMaker Ground Truth
Show answer & explanationAnswer & explanation
Correct answer: A. SageMaker Inference Container
The SageMaker Inference Container is responsible for loading the model artifact, running the inference code, and serving predictions when invoked through the real-time endpoint.
Why the other options are wrong
- B. SageMaker Processing Jobs are for data preprocessing, post-processing, and model evaluation, not for real-time inference serving.
- C. SageMaker Feature Store is for managing and serving features for training and inference, not for model serving itself.
- D. SageMaker Ground Truth is a data labeling service, unrelated to model deployment or inference serving.
SageMaker Inference Container
A Docker container orchestrated by Amazon SageMaker to host and serve machine learning models for real-time or batch inference.
- Contains model serving code (e.g., Flask, Gunicorn)
- Loads model artifacts from S3
- Executes prediction logic when invoked
- Can be custom or built-in SageMaker images
Memory trick: Containers Run Models Efficiently.