AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy

A data engineering team is building an MLOps pipeline for a fraud detection model. They need to ensure that all data used for training and inference is traceable and auditable. Specifically, they want to log all input features, model predictions, and any associated metadata for every inference request. This data will be used for model debugging, bias detection, and compliance auditing. Which SageMaker feature should be enabled on the real-time endpoint?

  1. ASageMaker Clarify.
  2. BSageMaker Data Capture.
  3. CSageMaker Model Monitor.
  4. DSageMaker Feature Store.
Show answer & explanation

Correct answer: B. SageMaker Data Capture.

SageMaker Data Capture is specifically designed to log inference requests and responses, including input features, model predictions, and custom metadata, to an S3 bucket. This captured data is crucial for traceability, auditing, debugging, and subsequent analysis by services like Model Monitor or Clarify.

Why the other options are wrong

  • A. SageMaker Clarify is used for bias detection and explainability, often *using* captured data, but it doesn't perform the data capture itself.
  • C. SageMaker Model Monitor uses captured data to detect drift but doesn't perform the initial data logging for auditing purposes.
  • D. SageMaker Feature Store manages and serves features, but it doesn't log the live inference requests and responses for auditing.

SageMaker Data Capture

A SageMaker capability that logs inference requests and responses, along with custom metadata, from real-time endpoints to an S3 bucket.

  • Provides raw data for auditing, debugging, and post-inference analysis.
  • Can capture both request payload and response payload.
  • Integrates with other SageMaker services like Model Monitor and Clarify.

Memory trick: Capture the data, for audit's sake, no mistakes!

More Machine Learning Implementation and Operations questions