A data science team is deploying a new object detection model to an Amazon SageMaker endpoint. This model is computationally intensive, and the team observes that inference requests sometimes time out under high load, even with sufficient instance scaling. They suspect that the default health check timeout might be too short for the model's initialization or a single inference request. Which SageMaker endpoint configuration parameter should they adjust to allow more time for the container to become healthy or process a request?
- AContainerStartupHealthCheckTimeoutInSeconds
- BInferenceLatencyThresholdInMilliseconds
- CModelDataDownloadTimeoutInSeconds
- DEndpointCreationTimeoutInSeconds
Show answer & explanationAnswer & explanation
Correct answer: A. ContainerStartupHealthCheckTimeoutInSeconds
The `ContainerStartupHealthCheckTimeoutInSeconds` parameter defines the maximum time that SageMaker waits for a container to respond to health checks during startup and also sets the maximum duration for a single inference request. If the model initialization is slow or a single inference request takes longer than the default (typically 60 seconds), increasing this value can prevent timeouts under load or during deployment.
Why the other options are wrong
- B. `InferenceLatencyThresholdInMilliseconds` is a metric for monitoring latency, not a configurable parameter to prevent timeouts.
- C. `ModelDataDownloadTimeoutInSeconds` controls the time allowed for downloading model artifacts, not for container startup or inference request processing.
- D. `EndpointCreationTimeoutInSeconds` defines the total time allowed for the entire endpoint creation process, not for individual container health or inference requests.
ContainerStartupHealthCheckTimeoutInSeconds
A SageMaker endpoint configuration parameter that specifies the maximum duration (in seconds) for a container to respond to health checks during startup, and also serves as the maximum timeout for a single inference request.
- Impacts both container startup and individual inference request timeouts.
- Default value is typically 60 seconds.
- Increasing it can prevent timeouts for large models or complex inferences.
- Configured at the endpoint configuration level.
Memory trick: Health check timeout: The container's patience for starting and responding to one request.