AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard

A media company uses a content moderation model deployed on a SageMaker real-time endpoint. The model's predictions are critical, and any downtime or service degradation could lead to significant business impact. The MLOps team wants to establish a robust CI/CD pipeline that automatically builds, tests, and deploys new model versions. A key requirement is to ensure that the new model is gradually rolled out to production, starting with a very small percentage of live traffic, and automatically rolls back if performance metrics (e.g., latency, error rate) exceed predefined thresholds. Which SageMaker deployment strategy, integrated with a CI/CD pipeline, should be used?

  1. ACanary Deployment
  2. BBlue/Green Deployment
  3. CA/B Testing
  4. DShadow Deployment
Show answer & explanation

Correct answer: A. Canary Deployment

Canary deployment allows a new model to be gradually rolled out to a small percentage of live traffic. With SageMaker's integration, this can be automated to monitor performance metrics and automatically roll back if thresholds are breached, ensuring minimal impact and safety for critical applications.

Why the other options are wrong

  • B. Blue/Green deployment involves switching all (or a large portion) of traffic at once or gradually, but often implies a full cutover rather than a cautious, small-percentage rollout with automatic rollback based on metrics.
  • C. A/B testing is for comparing two models' performance for business metrics, not primarily for a safe, gradual rollout with automatic rollback based on operational performance thresholds.
  • D. Shadow deployment processes a copy of traffic without impacting user responses, primarily for observation, not for gradually serving a new model to real users and automatically rolling back based on live metrics.

Canary Deployment (ML)

A deployment strategy where a new model version is gradually introduced to a small subset of live users/traffic, monitored for performance, and automatically rolled back if issues are detected, minimizing risk.

  • Gradual rollout to a small percentage of live traffic.
  • Continuous monitoring of key performance metrics.
  • Automated rollback if performance thresholds are breached.

Memory trick: Canaries sing a warning before full deployment.

More Machine Learning Implementation and Operations questions