AWS Certified Machine Learning – SpecialtyMachine Learning Implementation and OperationsMedium

A data science team is developing a new credit risk assessment model using Amazon SageMaker. They want to automate the entire machine learning workflow, from data preparation and model training to deployment and monitoring, using AWS services. New model versions should be automatically built and deployed whenever new training data becomes available or model code changes. Which AWS service combination provides the most suitable foundation for building such a CI/CD pipeline for ML models?

  1. AAWS S3, AWS Glue, Amazon Athena, Amazon SageMaker Studio
  2. BAWS Step Functions, AWS Batch, Amazon CloudWatch, Amazon SNS
  3. CAWS CodeCommit, AWS CodeBuild, AWS CodeDeploy, AWS Lambda
  4. DAWS CodeCommit, AWS CodeBuild, AWS CodePipeline, Amazon SageMaker Pipelines
Show answer & explanation

Correct answer: D. AWS CodeCommit, AWS CodeBuild, AWS CodePipeline, Amazon SageMaker Pipelines

AWS CodeCommit for source control, AWS CodeBuild for building and packaging, AWS CodePipeline for orchestrating the overall CI/CD workflow, and Amazon SageMaker Pipelines for managing the end-to-end ML workflow (data processing, training, deployment) are the ideal combination for an automated MLOps CI/CD pipeline.

Why the other options are wrong

  • A. S3, Glue, and Athena are for data storage and processing, while SageMaker Studio is an IDE. They are components *within* an ML workflow, not the CI/CD orchestrators.
  • B. Step Functions can orchestrate workflows, but SageMaker Pipelines is purpose-built for ML. Batch is for batch computing, CloudWatch and SNS for monitoring/notifications, not core CI/CD pipeline services.
  • C. AWS Lambda is for serverless functions, not the primary orchestrator for complex ML pipelines. CodeDeploy is more for traditional application deployments, less specific to ML models.

SageMaker MLOps CI/CD

An automated workflow for machine learning, typically using AWS Code services and SageMaker Pipelines, to manage the entire lifecycle from data to deployment.

  • Automates model building, testing, and deployment
  • Ensures reproducibility and traceability
  • Integrates with source control and monitoring
  • SageMaker Pipelines is key for ML-specific steps

Memory trick: Code's Pipeline Builds and Deploys SageMaker's ML.

More Machine Learning Implementation and Operations questions