AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard
A data scientist is deploying a new object detection model to an Amazon SageMaker endpoint. During initial testing, the endpoint occasionally fails to become `InService`, reporting `ClientError: Could not find model data` in the logs, even though the model artifact is correctly uploaded to S3. The data scientist suspects the model container is not fully ready to serve predictions before SageMaker considers it healthy. Which parameter should the data scientist adjust to resolve this issue?
- AContainerStartupHealthCheckTimeoutInSeconds
- BProductionVariantInitialInstanceCount
- CInstanceType
- DModelDataUrl
Show answer & explanationAnswer & explanation
Correct answer: A. ContainerStartupHealthCheckTimeoutInSeconds
The `ContainerStartupHealthCheckTimeoutInSeconds` parameter defines how long SageMaker waits for the container to pass its health checks during startup. If the model artifact is large or the container takes a long time to load the model and become ready, increasing this timeout allows SageMaker to wait longer before marking the container as unhealthy, resolving the `Could not find model data` issue when it's just a timing problem.
Why the other options are wrong
- B. ProductionVariantInitialInstanceCount sets the number of instances; it doesn't resolve issues with a single instance failing to start due to model loading times.
- C. InstanceType determines the compute resources; while it can affect loading time, adjusting the timeout directly addresses the 'not ready before health check' problem.
- D. ModelDataUrl specifies the S3 location of the model artifact; while crucial, it's not the issue if the artifact is 'correctly uploaded'.
ContainerStartupHealthCheckTimeoutInSeconds
A SageMaker endpoint configuration parameter that defines the maximum time (in seconds) that SageMaker waits for an inference container to pass its health checks during startup.
- Prevents premature health check failures for slow-starting containers.
- Useful for large models or complex container initialization.
- Default value is 60 seconds.
Memory trick: Timeout too short, model won't start.