AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard
A large e-commerce platform uses an ML model for real-time fraud detection. The model processes a massive volume of transactions, and the ML team needs to ensure that all input requests and corresponding model predictions are systematically recorded for auditing, debugging, and future model retraining purposes. This data must be stored securely and be easily accessible for analysis. Which SageMaker feature should they enable on their real-time endpoint?
- ASageMaker Model Monitor
- BSageMaker Batch Transform
- CSageMaker Clarify
- DSageMaker Data Capture
Show answer & explanationAnswer & explanation
Correct answer: D. SageMaker Data Capture
SageMaker Data Capture automatically records inference requests and responses for real-time endpoints and stores them in a specified Amazon S3 bucket. This feature is specifically designed for auditing, debugging, and collecting data for future model retraining, directly meeting the requirement for systematic and secure recording of inference data.
Why the other options are wrong
- A. SageMaker Model Monitor uses captured data to detect data and model quality drift, but Data Capture is the underlying mechanism for recording the data itself.
- B. SageMaker Batch Transform is for offline, batch inference, not for capturing real-time endpoint traffic.
- C. SageMaker Clarify is for detecting bias and explainability in models, not for capturing raw inference data.
SageMaker Data Capture
A feature of Amazon SageMaker real-time endpoints that automatically records inference requests and responses to an S3 bucket for auditing, debugging, and model retraining.
- Records input payloads and model predictions.
- Data stored securely in S3.
- Essential for monitoring, debugging, and drift detection baselines.
Memory trick: Capture data, capture insights.