AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsEasy

A machine learning team is optimizing their inference costs for a high-volume, low-latency application using SageMaker real-time endpoints. They have identified that the current instance type is underutilized during off-peak hours but experiences bottlenecks during peak times. They want to automatically adjust the number of instances based on the incoming traffic load to improve cost-efficiency and maintain performance. Which SageMaker feature should they implement?

  1. ASageMaker Data Wrangler
  2. BSageMaker Auto Scaling
  3. CSageMaker Batch Transform
  4. DSageMaker Model Monitor
Show answer & explanation

Correct answer: B. SageMaker Auto Scaling

SageMaker Auto Scaling automatically adjusts the number of instances provisioned for a real-time endpoint based on predefined metrics (like CPU utilization or invocation count) and scaling policies. This directly addresses the need to optimize cost and performance by adapting to varying traffic loads.

Why the other options are wrong

  • A. SageMaker Data Wrangler is for data preparation and feature engineering, unrelated to endpoint scaling.
  • C. SageMaker Batch Transform is for asynchronous processing of large datasets, not for real-time endpoint scaling.
  • D. SageMaker Model Monitor tracks model quality and data drift, not for scaling endpoint instances.

SageMaker Auto Scaling

A feature that automatically adjusts the number of instances for SageMaker real-time endpoints based on demand, optimizing cost and maintaining performance.

  • Uses CloudWatch metrics (e.g., CPU utilization, invocation count)
  • Scales out (adds instances) during peak load
  • Scales in (removes instances) during off-peak load
  • Helps manage costs and ensure low latency

Memory trick: Auto Scales to Save and Serve.

More Machine Learning Implementation and Operations questions