Step2Study
IT & TechnologyMLS-C01100% Free

AWS Certified Machine Learning – Specialty

Practice bank
243 Qs
Real exam
65 Qs
Time limit
170 min
Passing
Scaled score of 750 out of 1,000

Exam blueprint

Data Engineering
20%
Exploratory Data Analysis
20%
Modeling
30%
Machine Learning Implementation and Operations
30%

Practice

Untimed · instant feedback · 4 practice tests of 90 questions

Questions per test

Custom practice

Flashcard on every question Mental map when you miss

Exam simulation

4 timed tests · 90 questions each · 235 min · pass 75% · 243 questions in the bank

+50 XP per test · +100 XP for a pass

Random simulation (weighted by domain)

Everything is open to everyone. Create a free account to save scores, XP, badges and get progress emails.

Free study resources

All resources →

Study with friends

Challenge a friend to beat your score.

AWS Certified Machine Learning – Specialty practice test questions

Sample questions from the 243-question bank, with answers and explanations.

All questions
  1. 1. A company is using a deep learning model for image classification deployed on an Amazon SageMaker endpoint. To reduce inference costs and improve throughput, they want to optimize the model for specific hardware accelerators available on SageMaker instances. The optimization process needs to be framework-agnostic and produce a deployable artifact. Which SageMaker capability should they leverage?

    Machine Learning Implementation and Operations

    • A. SageMaker Distributed Training
    • B. SageMaker Inference Recommender
    • C. SageMaker Training Compiler
    • D. SageMaker Neo
    Show answer

    D. SageMaker Neo

    SageMaker Neo is a model compilation service that optimizes models from various frameworks for specific hardware platforms (including SageMaker instances with accelerators) to achieve faster inference and lower costs. It produces a compiled, deployable artifact.

  2. 2. A pharmaceutical company is developing a machine learning model to assist in drug discovery. The model is computationally intensive, requiring significant GPU resources for inference, but the inference requests are infrequent and latency is not extremely critical (a few seconds is acceptable). The company wants to optimize costs for these inference workloads. Which SageMaker feature should they consider to achieve cost-effective GPU-powered inference?

    Machine Learning Implementation and Operations

    • A. SageMaker Multi-Model Endpoints
    • B. SageMaker Elastic Inference
    • C. SageMaker Serverless Inference
    • D. SageMaker Batch Transform
    Show answer

    B. SageMaker Elastic Inference

    SageMaker Elastic Inference allows you to attach GPU acceleration to CPU-based EC2 instances or SageMaker instances at a fraction of the cost of a full GPU instance. Since the model requires GPU resources but inference is infrequent and latency is not extremely critical, Elastic Inference provides a cost-effective way to get GPU acceleration without paying for oversized GPU instances, aligning perfectly with the cost optimization goal for infrequent workloads.

  3. 3. A data science team is deploying a new object detection model to an Amazon SageMaker endpoint. During initial testing, they observe that the endpoint instance takes longer than expected to become 'InService' after deployment or scaling events, occasionally leading to failed deployments. The model's container initializes several large dependencies at startup. Which SageMaker endpoint configuration parameter should they adjust to allow more time for the container to fully initialize and pass health checks?

    Machine Learning Implementation and Operations

    • A. ContainerStartupHealthCheckTimeoutInSeconds
    • B. ProductionVariant.InstanceType
    • C. ProductionVariant.InitialInstanceCount
    • D. ModelDataDownloadTimeoutInSeconds
    Show answer

    A. ContainerStartupHealthCheckTimeoutInSeconds

    The `ContainerStartupHealthCheckTimeoutInSeconds` parameter defines the maximum amount of time SageMaker waits for the container to emit a successful response to the `/ping` health check. If a container has large dependencies or complex initialization logic, increasing this timeout allows more time for it to become ready, preventing premature health check failures and deployment issues.

  4. 4. A large e-commerce platform uses an ML model for real-time fraud detection. The model processes millions of transactions per hour, and high availability is paramount. Any downtime or performance degradation could lead to significant financial losses. The team needs to deploy the model in a way that provides automatic failover and resilience against Availability Zone (AZ) outages. Which SageMaker deployment option should they choose?

    Machine Learning Implementation and Operations

    • A. Deploy to a SageMaker Multi-AZ endpoint.
    • B. Deploy multiple single-AZ endpoints and manage failover with Route 53.
    • C. Deploy to a SageMaker serverless endpoint.
    • D. Deploy to a single SageMaker endpoint with auto-scaling enabled.
    Show answer

    A. Deploy to a SageMaker Multi-AZ endpoint.

    A SageMaker Multi-AZ endpoint automatically distributes inference requests across instances in multiple Availability Zones. If an AZ becomes unhealthy, traffic is automatically routed to healthy instances in other AZs, providing built-in high availability and resilience against AZ outages without manual intervention or additional services like Route 53 for failover management.

  5. 5. A startup is building a personalized content recommendation system. They need to frequently update the recommendations based on user interactions and content changes, requiring daily model retraining. The entire MLOps workflow, from data ingestion and processing to model training, evaluation, and deployment, needs to be fully automated and orchestrated. Which AWS service is best suited for defining and orchestrating this end-to-end machine learning workflow?

    Machine Learning Implementation and Operations

    • A. Amazon SageMaker Pipelines
    • B. AWS Step Functions
    • C. AWS CodeBuild
    • D. AWS Lambda
    Show answer

    A. Amazon SageMaker Pipelines

    Amazon SageMaker Pipelines is specifically designed for building, automating, and managing end-to-end machine learning workflows (MLOps pipelines), making it the best choice for orchestrating daily model retraining, evaluation, and deployment.

  6. 6. A media company uses a content moderation model deployed on a SageMaker real-time endpoint. They want to introduce a new, improved version of the model to a small, controlled group of users (e.g., 5%) to observe its real-world performance and stability before a full rollout. They need to be able to quickly revert to the old model if any major issues arise, minimizing impact on the majority of users. Which deployment strategy should they employ?

    Machine Learning Implementation and Operations

    • A. Canary deployment
    • B. Shadow deployment
    • C. In-place update
    • D. Blue/Green deployment
    Show answer

    A. Canary deployment

    Canary deployment involves gradually rolling out a new model version to a small subset of users (the 'canary' group) while the majority of traffic still goes to the old version. This allows for real-world testing and monitoring of the new model's performance and stability with minimal risk, as issues affect only a small percentage of users, and a quick rollback is possible by simply removing the canary variant.

  7. 7. A large e-commerce platform uses a recommendation engine powered by a machine learning model deployed on Amazon SageMaker. During peak shopping seasons, the inference requests can surge unexpectedly, leading to increased latency and potential service disruptions. The platform needs a solution that can automatically adjust the endpoint's capacity to handle these unpredictable traffic spikes without manual intervention and ensure consistent low latency. Which combination of SageMaker features should be implemented?

    Machine Learning Implementation and Operations

    • A. SageMaker Model Monitor and SageMaker Neo.
    • B. SageMaker Endpoint Auto Scaling and Multi-AZ deployment.
    • C. SageMaker Data Wrangler and SageMaker Pipelines.
    • D. SageMaker Batch Transform and Elastic Inference.
    Show answer

    B. SageMaker Endpoint Auto Scaling and Multi-AZ deployment.

    SageMaker Endpoint Auto Scaling dynamically adjusts the number of inference instances based on load, ensuring capacity during traffic spikes. Multi-AZ deployment distributes these instances across multiple Availability Zones, providing high availability and resilience against single-point failures, which is crucial for consistent low latency and service continuity.

  8. 8. A healthcare provider is deploying a diagnostic ML model that processes sensitive patient data. Regulatory compliance (e.g., HIPAA) requires strict data provenance and an immutable audit trail of every inference request, including the input data, model version, and prediction results. This audit trail is essential for post-hoc analysis, debugging, and demonstrating compliance. Which SageMaker feature, combined with appropriate AWS logging, can best meet these requirements?

    Machine Learning Implementation and Operations

    • A. SageMaker Pipelines with model versioning.
    • B. SageMaker Asynchronous Inference with data capture.
    • C. SageMaker Real-Time Endpoint with data capture enabled.
    • D. SageMaker Model Monitor with custom metrics.
    Show answer

    C. SageMaker Real-Time Endpoint with data capture enabled.

    SageMaker Real-Time Endpoints can be configured with data capture, which automatically saves a configurable percentage of inference requests and responses (including input data and prediction results) to an S3 bucket. This raw data, combined with CloudWatch logs for model versioning and timestamps, creates an immutable audit trail necessary for regulatory compliance and deep analysis.

  9. 9. An e-commerce company uses an ML model for real-time product recommendations. The model processes millions of inference requests per hour. Due to the high volume and the need for auditing and debugging, the company requires that all inference requests and responses, including the input and output data, be captured and stored for future analysis. Which SageMaker feature enables this capability?

    Machine Learning Implementation and Operations

    • A. SageMaker Data Capture
    • B. SageMaker Debugger
    • C. SageMaker Clarify
    • D. SageMaker Model Monitor
    Show answer

    A. SageMaker Data Capture

    SageMaker Data Capture allows you to configure a real-time inference endpoint to automatically capture and save input and output data, as well as model predictions, to an Amazon S3 bucket. This data can then be used for model monitoring, debugging, auditing, and retraining purposes.

  10. 10. A financial institution is developing a new credit risk assessment model. They need a centralized, versioned, and low-latency storage solution for features that can be used consistently across both model training and real-time inference. The solution must support point-in-time correctness for historical training data and fast retrieval for online predictions. Which SageMaker service is designed to meet these requirements?

    Machine Learning Implementation and Operations

    • A. Amazon S3 for raw data storage.
    • B. SageMaker Feature Store.
    • C. SageMaker Model Registry for model versioning.
    • D. Amazon DynamoDB for feature storage.
    Show answer

    B. SageMaker Feature Store.

    SageMaker Feature Store is specifically designed to address the challenges of feature management for ML, providing a centralized repository for features. It supports both online (low-latency retrieval for inference) and offline (point-in-time correctness for training) access, along with versioning and consistent feature definitions.

  11. 11. A machine learning engineer needs to deploy a new model that accurately predicts customer churn. Before fully replacing the existing model, the engineer wants to compare the performance of the new model against the current production model using real-time traffic, but only for a small percentage of users, without impacting the majority. Which SageMaker deployment strategy should the engineer use?

    Machine Learning Implementation and Operations

    • A. Blue/Green Deployment
    • B. Batch Transform
    • C. Shadow Deployment
    • D. A/B Testing
    Show answer

    D. A/B Testing

    A/B testing on SageMaker endpoints allows routing a small percentage of live traffic to a new model (variant) while the majority of traffic goes to the current model (production), enabling direct comparison of performance metrics with minimal user impact.

  12. 12. A data science team has developed a new fraud detection model using Amazon SageMaker. The team wants to deploy this model to a production environment where it needs to handle real-time inference requests with low latency and high availability. The model is relatively small and can be loaded into memory quickly. Which deployment option is most suitable for this scenario?

    Machine Learning Implementation and Operations

    • A. SageMaker Neo compilation
    • B. Batch transform job
    • C. Asynchronous inference endpoint
    • D. SageMaker endpoint
    Show answer

    D. SageMaker endpoint

    SageMaker endpoints are designed for real-time, low-latency, and high-availability inference, making them ideal for interactive applications like fraud detection.

  13. 13. A data science team is developing a critical machine learning model for real-time anomaly detection. They need to ensure that the model is always available and highly resilient to instance failures or Availability Zone (AZ) outages. The application requires very low latency for predictions. Which SageMaker endpoint configuration best addresses these requirements?

    Machine Learning Implementation and Operations

    • A. Deploying the model to a multi-instance SageMaker endpoint within a single Availability Zone.
    • B. Deploying the model to a single-instance SageMaker endpoint in a single Availability Zone.
    • C. Deploying the model to a SageMaker serverless endpoint for automatic scaling.
    • D. Deploying the model to a multi-instance SageMaker endpoint across multiple Availability Zones.
    Show answer

    D. Deploying the model to a multi-instance SageMaker endpoint across multiple Availability Zones.

    Deploying a multi-instance SageMaker endpoint across multiple Availability Zones provides the highest level of availability and resilience. If one instance or an entire AZ fails, traffic is automatically routed to healthy instances in other AZs, ensuring continuous operation and low latency.

  14. 14. A machine learning engineer is deploying a real-time inference endpoint for a critical fraud detection model using Amazon SageMaker. The model requires high availability and needs to be resilient to instance failures within a single Availability Zone. Which configuration should the engineer primarily focus on to address this requirement?

    Machine Learning Implementation and Operations

    • A. Implementing Elastic Inference
    • B. Configuring a Multi-AZ endpoint
    • C. Enabling SageMaker Model Monitor
    • D. Setting up SageMaker Pipelines
    Show answer

    B. Configuring a Multi-AZ endpoint

    Configuring a Multi-AZ endpoint for Amazon SageMaker ensures that the inference endpoint is deployed across multiple Availability Zones. This provides high availability and resilience against single Availability Zone outages or instance failures within a single AZ.

  15. 15. A retail company is deploying a new personalized recommendation model. To minimize risk, they want to gradually shift traffic from the old model (Model A) to the new model (Model B), carefully monitoring Model B's performance and impact on key business metrics like conversion rate. If Model B performs worse than expected, they need to be able to quickly revert to Model A. Which deployment strategy should they employ?

    Machine Learning Implementation and Operations

    • A. Shadow Deployment
    • B. A/B Testing
    • C. Canary Deployment
    • D. Blue/Green Deployment
    Show answer

    C. Canary Deployment

    Canary deployment allows gradually routing a small percentage of live traffic to the new model (Model B) while the majority still goes to the old model (Model A). This enables real-time monitoring of the new model's performance with actual user traffic. If issues arise, traffic can be instantly rolled back to Model A, minimizing impact. This fits the requirement for gradual shifting and quick rollback.

  16. 16. A data scientist needs to deploy a PyTorch model to an AWS IoT Greengrass device for local inference. The device has limited computational resources and memory. Before deployment, the model needs to be optimized for performance and size on the target hardware. Which SageMaker feature should the data scientist use to achieve this?

    Machine Learning Implementation and Operations

    • A. SageMaker Neo
    • B. SageMaker Clarify
    • C. SageMaker Model Monitor
    • D. SageMaker Pipelines
    Show answer

    A. SageMaker Neo

    SageMaker Neo is designed to compile machine learning models to optimize them for specific hardware targets (like IoT Greengrass devices), resulting in up to 2x faster inference and reduced memory footprint.

  17. 17. A data science team is building a large language model (LLM) for a natural language processing application. The model is too large to fit into the memory of a single GPU instance or to process efficiently on a single instance during inference. The team needs to deploy this LLM on Amazon SageMaker, optimizing for memory utilization and inference latency across multiple instances. Which SageMaker feature is specifically designed to handle the deployment of very large models by distributing them across multiple devices or instances?

    Machine Learning Implementation and Operations

    • A. SageMaker Neo
    • B. SageMaker Model Parallelism (Inference)
    • C. SageMaker Elastic Inference
    • D. SageMaker Multi-Model Endpoints
    Show answer

    B. SageMaker Model Parallelism (Inference)

    SageMaker Model Parallelism (Inference) is specifically designed to deploy very large models, like LLMs, by automatically splitting the model across multiple GPU instances or devices to overcome memory limitations and improve inference latency.

  18. 18. A data science team is developing a new credit risk assessment model. They need a centralized, versioned repository to store, manage, and share features across different machine learning models and teams. This repository must support both online inference and offline training. Which AWS service is best suited for this requirement?

    Machine Learning Implementation and Operations

    • A. Amazon RDS
    • B. Amazon SageMaker Feature Store
    • C. Amazon DynamoDB
    • D. Amazon S3
    Show answer

    B. Amazon SageMaker Feature Store

    Amazon SageMaker Feature Store is a purpose-built service for machine learning that provides a centralized repository for features. It supports both online (low-latency) and offline (batch) access for training and inference, and includes capabilities for feature versioning and sharing, directly meeting all stated requirements.

  19. 19. A data science team is developing a fraud detection model. They need a centralized, version-controlled repository to store, share, and manage features used across multiple models and teams, ensuring consistency between training and inference. The solution must support both online (low-latency) and offline (batch) access patterns. Which AWS service is best suited for this requirement?

    Machine Learning Implementation and Operations

    • A. Amazon SageMaker Feature Store
    • B. Amazon S3
    • C. AWS Glue Data Catalog
    • D. Amazon DynamoDB
    Show answer

    A. Amazon SageMaker Feature Store

    Amazon SageMaker Feature Store is a purpose-built service for storing, updating, and serving machine learning features for training and inference. It supports both online (low-latency) and offline (batch processing) access, ensuring consistency and reusability.

  20. 20. A data engineering team is building a new feature store to manage features for various machine learning models across their organization. They need a solution that can serve features with very low latency for real-time inference, as well as provide historical feature data for model training and batch inference. Which SageMaker component or mode of operation of the Feature Store meets both of these requirements?

    Machine Learning Implementation and Operations

    • A. SageMaker Feature Store - Online Store only
    • B. SageMaker Feature Store - Offline Store only
    • C. SageMaker Feature Store - Online and Offline Stores
    • D. SageMaker Feature Store - Data Catalog integration
    Show answer

    C. SageMaker Feature Store - Online and Offline Stores

    Amazon SageMaker Feature Store is designed with both an Online Store and an Offline Store. The Online Store provides low-latency access to the latest feature values for real-time inference, while the Offline Store stores historical feature data for model training, batch inference, and analytical purposes. Using both modes simultaneously ensures all requirements are met.

  21. 21. A data science team is developing a new credit scoring model. They want to ensure that the model is fair and unbiased across different demographic groups (e.g., age, gender, ethnicity) and that its decisions are explainable to stakeholders and regulatory bodies. Which Amazon SageMaker feature should they integrate into their MLOps workflow to address these requirements?

    Machine Learning Implementation and Operations

    • A. Amazon SageMaker Model Monitor
    • B. Amazon SageMaker Clarify
    • C. Amazon SageMaker Neo
    • D. Amazon SageMaker Feature Store
    Show answer

    B. Amazon SageMaker Clarify

    Amazon SageMaker Clarify is specifically designed to help machine learning developers detect potential bias in their data and models, and to provide explainability for model predictions. It supports various bias metrics and explainability techniques (like SHAP and LIME) to ensure fairness and transparency, which are critical for sensitive applications like credit scoring and regulatory compliance.

  22. 22. A manufacturing company uses a machine learning model for predictive maintenance on factory equipment. The model is deployed on SageMaker. Over time, the types of equipment failures and their underlying causes evolve due to new machinery, production processes, and environmental factors. The current model, trained on historical data, fails to accurately predict these new failure modes, leading to increased downtime. This phenomenon is known as concept drift. What is the most effective MLOps strategy to address this specific challenge?

    Machine Learning Implementation and Operations

    • A. Implement A/B testing with a different model architecture.
    • B. Set up SageMaker Model Monitor to detect data drift and trigger retraining.
    • C. Regularly retrain the model on periodically collected new data.
    • D. Use SageMaker Clarify to analyze model bias and explainability.
    Show answer

    C. Regularly retrain the model on periodically collected new data.

    Concept drift occurs when the relationship between input features and the target variable changes over time. While Model Monitor can detect drift, the most direct and effective strategy to address concept drift is to regularly retrain the model on the most recently collected data that reflects the new underlying concepts. This ensures the model learns the current patterns and relationships.

  23. 23. A data science team needs to implement a robust MLOps practice where new model versions are automatically trained, evaluated, and deployed based on new data or code changes. They want to define a series of steps for this entire workflow, including data preprocessing, training, model evaluation, and conditional deployment, all orchestrated within SageMaker. Which AWS service is purpose-built for defining and managing such end-to-end ML workflows?

    Machine Learning Implementation and Operations

    • A. Amazon SageMaker Studio
    • B. Amazon SageMaker Pipelines
    • C. Amazon EventBridge
    • D. AWS CodePipeline
    Show answer

    B. Amazon SageMaker Pipelines

    Amazon SageMaker Pipelines is a purpose-built MLOps service for building, automating, and managing end-to-end machine learning workflows. It allows data scientists and ML engineers to define a sequence of steps (like data processing, training, evaluation, registration, and conditional deployment), track lineage, and automate the entire ML lifecycle, directly addressing the need for a robust, orchestrated MLOps practice within SageMaker.

  24. 24. A data science team is deploying a new object detection model to an Amazon SageMaker endpoint. This model is computationally intensive, and the team observes that inference requests sometimes time out under high load, even with sufficient instance scaling. They suspect that the default health check timeout might be too short for the model's initialization or a single inference request. Which SageMaker endpoint configuration parameter should they adjust to allow more time for the container to become healthy or process a request?

    Machine Learning Implementation and Operations

    • A. ContainerStartupHealthCheckTimeoutInSeconds
    • B. InferenceLatencyThresholdInMilliseconds
    • C. ModelDataDownloadTimeoutInSeconds
    • D. EndpointCreationTimeoutInSeconds
    Show answer

    A. ContainerStartupHealthCheckTimeoutInSeconds

    The `ContainerStartupHealthCheckTimeoutInSeconds` parameter defines the maximum time that SageMaker waits for a container to respond to health checks during startup and also sets the maximum duration for a single inference request. If the model initialization is slow or a single inference request takes longer than the default (typically 60 seconds), increasing this value can prevent timeouts under load or during deployment.

  25. 25. A financial services company is developing a new credit risk assessment model. They need to ensure that the model's predictions are fair and unbiased across different demographic groups (e.g., age, gender) and that the model's decision-making process is transparent and explainable to comply with regulatory requirements. Which Amazon SageMaker feature should the data science team use to achieve these goals during model development and post-deployment monitoring?

    Machine Learning Implementation and Operations

    • A. SageMaker Feature Store for feature consistency.
    • B. SageMaker Model Monitor for data drift detection.
    • C. SageMaker Clarify for bias detection and explainability.
    • D. SageMaker Pipelines for MLOps automation.
    Show answer

    C. SageMaker Clarify for bias detection and explainability.

    SageMaker Clarify is specifically designed to detect potential bias in ML models and provide explainability for their predictions, directly addressing the requirements for fairness and transparency across demographic groups and regulatory compliance.

AWS Certified Machine Learning – Specialty flashcards

Tap a card to flip it. 144 flashcards in the full deck.

  • SageMaker Neo for Inference

    Flip card

    A SageMaker service that compiles ML models from various frameworks into an optimized executable for specific hardware targets, enhancing inference performance and efficiency.

    • Reduces inference latency and cost.
    • Supports various frameworks (TensorFlow, PyTorch, MXNet).
    • Optimizes for cloud instances (e.g., with GPUs) and edge devices.
    Study this card →
  • SageMaker Elastic Inference

    Flip card

    A feature that allows attaching GPU-powered inference acceleration to Amazon EC2 and SageMaker instances, reducing the cost of deep learning inference.

    • Provides GPU acceleration at a lower cost than full GPU instances.
    • Suitable for models requiring GPU but don't fully utilize a dedicated GPU.
    • Integrates with existing CPU instances.
    Study this card →
  • ContainerStartupHealthCheckTimeoutInSeconds

    Flip card

    A SageMaker endpoint configuration parameter that sets the maximum time (in seconds) allowed for a container to pass its initial health check (`/ping`) during deployment or scaling.

    • Crucial for containers with long initialization processes.
    • Prevents premature deployment failures.
    • Default value is usually 300 seconds (5 minutes).
    Study this card →
  • SageMaker Multi-AZ Endpoint

    Flip card

    A SageMaker deployment option that automatically distributes model instances across multiple AWS Availability Zones, providing high availability and automatic failover for real-time inference.

    • Ensures resilience against Availability Zone outages.
    • Automatically routes traffic to healthy instances.
    • Provides built-in high availability for critical models.
    Study this card →
  • SageMaker Pipelines

    Flip card

    Amazon SageMaker Pipelines is a purpose-built MLOps service for creating, automating, and managing end-to-end machine learning workflows, from data preparation to model deployment.

    • Orchestrates multi-step ML workflows.
    • Supports continuous integration and continuous delivery (CI/CD) for ML.
    • Provides lineage tracking and reproducibility for ML models.
    Study this card →
  • Canary Deployment (ML)

    Flip card

    A deployment strategy where a new version of an ML model is rolled out to a small, controlled percentage of live users, allowing for real-world testing and monitoring before a full rollout.

    • Minimizes risk by limiting exposure to new model.
    • Allows observation of performance with live traffic.
    • Facilitates quick rollback to the old version.
    Study this card →
  • SageMaker Auto Scaling & Multi-AZ

    Flip card

    SageMaker Endpoint Auto Scaling dynamically adjusts inference instance counts based on demand, while Multi-AZ deployment distributes instances across Availability Zones for high availability and fault tolerance.

    • Auto Scaling adjusts capacity for traffic changes.
    • Multi-AZ ensures high availability and resilience.
    • Together, they provide robust, scalable, and fault-tolerant inference.
    Study this card →
  • SageMaker Data Capture

    Flip card

    A feature for SageMaker inference endpoints (real-time or asynchronous) that automatically captures and saves input and output data for inference requests to S3.

    • Provides a raw, immutable record of inferences
    • Essential for auditing, debugging, and compliance
    • Captures request payload, response payload, and metadata
    Study this card →
  • SageMaker Feature Store

    Flip card

    A purpose-built service for MLOps that provides a centralized repository for storing, discovering, and sharing ML features for both training and inference.

    • Includes an online store for low-latency inference and an offline store for training.
    • Ensures feature consistency between training and inference (prevents skew).
    • Supports point-in-time lookups for historical feature values.
    Study this card →
  • A/B Testing (ML Deployment)

    Flip card

    A deployment strategy where different versions of a machine learning model are served to different segments of live user traffic to compare their performance metrics directly.

    • Routes a percentage of real-time traffic to new model.
    • Allows direct comparison of model performance.
    • Minimizes risk by limiting exposure of new model.
    Study this card →
  • SageMaker Endpoint

    Flip card

    A fully managed, continuously running HTTPS endpoint for real-time machine learning model inference.

    • Provides low-latency, high-throughput inference.
    • Supports auto-scaling to handle varying load.
    • Integrates with various SageMaker features like model monitoring.
    Study this card →
  • Canary Deployment

    Flip card

    A deployment strategy where a new version of an application or model is incrementally rolled out to a small subset of users before a full rollout.

    • Reduces risk by exposing new version to limited users
    • Allows real-time performance monitoring with live traffic
    • Enables quick rollback to the old version if issues arise
    Study this card →
  • SageMaker Neo

    Flip card

    A SageMaker feature that compiles machine learning models to optimize them for specific hardware platforms, improving inference performance and reducing memory footprint.

    • Supports various frameworks (TensorFlow, PyTorch, MXNet, etc.).
    • Optimizes for edge devices (IoT Greengrass) and cloud instances.
    • Results in faster inference and smaller model size.
    Study this card →
  • SageMaker Model Parallelism (Inference)

    Flip card

    A SageMaker capability that enables the deployment of very large deep learning models by distributing the model's layers or components across multiple GPU instances or devices for efficient inference.

    • Distributes large models across multiple GPUs/instances.
    • Overcomes single-device memory limitations.
    • Optimizes inference latency and throughput for LLMs.
    Study this card →
  • SageMaker Feature Store (Online/Offline)

    Flip card

    A fully managed service that provides a centralized repository for machine learning features, consisting of an Online Store for low-latency real-time inference and an Offline Store for historical data for training and batch inference.

    • Online Store for real-time, low-latency feature retrieval.
    • Offline Store for historical data, training, and batch inference.
    • Ensures consistency between training and inference features.
    Study this card →
  • Amazon SageMaker Clarify

    Flip card

    An Amazon SageMaker feature that provides tools to detect bias in machine learning datasets and models, and to help explain model predictions, promoting fairness and transparency.

    • Detects pre-training, post-training, and post-deployment bias.
    • Generates model explanations using techniques like SHAP and LIME.
    • Helps ensure regulatory compliance and ethical AI.
    Study this card →
  • Concept Drift Mitigation

    Flip card

    Strategies employed to maintain ML model performance when the underlying relationship between inputs and outputs changes over time.

    • Regular retraining with fresh data is key
    • Monitoring can detect drift, but retraining is the fix
    • Can involve adaptive learning techniques
    Study this card →
  • Amazon SageMaker Pipelines

    Flip card

    A purpose-built MLOps service within SageMaker for creating, automating, and managing end-to-end machine learning workflows, including data preparation, training, evaluation, and deployment.

    • Automates the entire ML lifecycle.
    • Provides lineage tracking for reproducibility.
    • Supports conditional logic for deployment decisions.
    Study this card →
  • SageMaker Clarify

    Flip card

    A SageMaker capability that helps detect bias in machine learning datasets and models, and provides tools to explain model predictions, promoting fairness and transparency.

    • Detects pre-training and post-training bias.
    • Generates model explanations (e.g., SHAP, LIME).
    • Supports monitoring for bias and explainability in production.
    Study this card →
  • SageMaker Inference Container

    Flip card

    A Docker container orchestrated by Amazon SageMaker to host and serve machine learning models for real-time or batch inference.

    • Contains model serving code (e.g., Flask, Gunicorn)
    • Loads model artifacts from S3
    • Executes prediction logic when invoked
    Study this card →
  • Conditional Model Retraining

    Flip card

    An MLOps practice where a model is retrained only when specific performance degradation or data drift conditions are met, rather than on a fixed schedule.

    • Reduces computational costs and resource usage associated with unnecessary retraining.
    • Ensures model quality by retraining only when genuinely needed.
    • Often implemented using monitoring tools like SageMaker Model Monitor to detect triggers.
    Study this card →
  • Blue/Green Deployment (ML)

    Flip card

    A deployment strategy where two identical production environments (Blue: current, Green: new) are maintained. Traffic is shifted from Blue to Green, allowing for quick rollback if issues occur.

    • Minimizes downtime during deployment.
    • Enables instant rollback to the previous stable version.
    • Requires double the infrastructure capacity temporarily.
    Study this card →
  • SageMaker Model Monitor

    Flip card

    A SageMaker service that continuously monitors the quality of ML models in production, detecting issues such as data drift, model drift, and data quality anomalies.

    • Compares real-time inference data against a baseline.
    • Can trigger alerts or automated retraining pipelines.
    • Helps maintain model performance over time.
    Study this card →
  • SageMaker Endpoint Encryption (In-transit)

    Flip card

    Ensuring data exchanged between a client and a SageMaker endpoint is encrypted while it's moving across a network.

    • Achieved primarily through HTTPS/TLS.
    • Protects sensitive inference requests and responses.
    • Standard practice for secure ML deployments.
    Study this card →

Questions are original practice items written to match the published exam objectives. Step2Study is not affiliated with or endorsed by any certification body.