AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy

A data science team has developed a new fraud detection model using Amazon SageMaker. The team wants to deploy this model to a production environment where it needs to handle real-time inference requests with low latency and high availability. The model is relatively small and can be loaded into memory quickly. Which deployment option is most suitable for this scenario?

  1. ASageMaker Neo compilation
  2. BBatch transform job
  3. CAsynchronous inference endpoint
  4. DSageMaker endpoint
Show answer & explanation

Correct answer: D. SageMaker endpoint

SageMaker endpoints are designed for real-time, low-latency, and high-availability inference, making them ideal for interactive applications like fraud detection.

Why the other options are wrong

  • A. SageMaker Neo optimizes models for specific hardware, but doesn't directly provide a real-time serving endpoint.
  • B. Batch transform is for offline processing of large datasets, not real-time inference.
  • C. Asynchronous inference is for requests with large payload sizes or long processing times, not typically for low-latency real-time fraud detection.

SageMaker Endpoint

A fully managed, continuously running HTTPS endpoint for real-time machine learning model inference.

  • Provides low-latency, high-throughput inference.
  • Supports auto-scaling to handle varying load.
  • Integrates with various SageMaker features like model monitoring.

Memory trick: Real-time needs a persistent door, not a batch truck.

More Machine Learning Implementation and Operations questions