AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsHard

A startup is building a personalized content recommendation system. They need to experiment with different model architectures and feature sets frequently to improve recommendation quality. Each experiment involves training a new model, deploying it, and evaluating its performance in a live environment against the current production model, often for a limited period to collect metrics. They want to ensure that these experiments can be run in parallel without affecting the core user experience for the majority of users. Which MLOps practice is best suited for this scenario?

  1. ACanary Deployment.
  2. BA/B Testing.
  3. CShadow Deployment.
  4. DBlue/Green Deployment.
Show answer & explanation

Correct answer: B. A/B Testing.

A/B testing is the most appropriate MLOps practice for this scenario. It allows the startup to compare the performance of different model architectures and feature sets (versions A and B) in a live environment by routing distinct, isolated groups of users to each version. This enables direct comparison of metrics and user experience impact for different models, without affecting the majority of users, and is designed for experimenting with distinct strategies over time.

Why the other options are wrong

  • A. Canary deployment is for gradually rolling out a single new model version to a small percentage of *all* users, primarily for risk mitigation during updates, not for comparing distinct models to separate user groups for experimentation.
  • C. Shadow deployment is for passive testing where the new model's predictions are not served to users, which doesn't fit the requirement of 'evaluating its performance in a live environment against the current production model' for distinct user groups.
  • D. Blue/Green deployment is for full-scale cutovers, not for controlled experimentation with distinct user groups.

A/B Testing (ML)

An experimental framework used to compare two or more versions of a machine learning model or system in a live production environment to determine which performs better.

  • Routes distinct user segments to different model versions (A vs. B).
  • Enables direct comparison of business and ML metrics for different strategies.
  • Ideal for iterative improvement and validating new model architectures/features.

Memory trick: Experiment and compare, let the users declare, which model's beyond compare!

More Machine Learning Implementation and Operations questions