AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy
A data science team has developed a new fraud detection model using Amazon SageMaker. The team wants to deploy this model to a production environment where it needs to handle real-time inference requests with low latency and high availability. The model is relatively small and can be loaded into memory quickly. Which deployment option is most suitable for this scenario?
- ASageMaker Neo compilation
- BBatch transform job
- CAsynchronous inference endpoint
- DSageMaker endpoint
Show answer & explanationAnswer & explanation
Correct answer: D. SageMaker endpoint
SageMaker endpoints are designed for real-time, low-latency, and high-availability inference, making them ideal for interactive applications like fraud detection.
Why the other options are wrong
- A. SageMaker Neo optimizes models for specific hardware, but doesn't directly provide a real-time serving endpoint.
- B. Batch transform is for offline processing of large datasets, not real-time inference.
- C. Asynchronous inference is for requests with large payload sizes or long processing times, not typically for low-latency real-time fraud detection.
SageMaker Endpoint
A fully managed, continuously running HTTPS endpoint for real-time machine learning model inference.
- Provides low-latency, high-throughput inference.
- Supports auto-scaling to handle varying load.
- Integrates with various SageMaker features like model monitoring.
Memory trick: Real-time needs a persistent door, not a batch truck.