AWS Certified Machine Learning – Specialty practice questions

243 free questions with answers and explanations.

Practice test
  1. 1.A retail company wants to test two different versions of a recommendation model (Model A and Model B) in a production environment to determine which one generates higher customer engagement before fully rolling out the superior model. They need to direct a small percentage (e.g., 10%) of live traffic to Model B and the remaining traffic to Model A, collect performance metrics for both, and compare the results. Which deployment strategy is best suited for this scenario?Machine Learning Implementation and Operations
  2. 2.A data science team has developed a new image classification model using a custom deep learning framework that is not natively supported by Amazon SageMaker's built-in algorithms or common framework containers (like TensorFlow or PyTorch). They need to deploy this model for real-time inference on SageMaker, ensuring that their unique framework and dependencies are correctly packaged and available at runtime. Which approach should the team take to deploy their model?Machine Learning Implementation and Operations
  3. 3.A data science team is developing a critical machine learning model for real-time anomaly detection. The model needs to be updated frequently, but the team wants to minimize downtime and risk during deployments. They also need a mechanism to quickly roll back to the previous stable version if issues are detected with the new model. Which deployment strategy is most suitable for this scenario?Machine Learning Implementation and Operations
  4. 4.An e-commerce company uses an ML model for real-time product recommendations. The model processes millions of inference requests daily, leading to high operational costs for the SageMaker endpoint instances. The data science team observes that the model's performance is rarely impacted by significant changes in the underlying data distribution, and retraining is only necessary occasionally (e.g., quarterly). They want to optimize costs by reducing the frequency of model retraining while ensuring adequate model performance over time. Which MLOps practice would be most beneficial in this scenario?Machine Learning Implementation and Operations
  5. 5.A data science team has developed a new machine learning model to predict customer churn. They need to deploy this model to production in a way that minimizes downtime and allows for a quick rollback if issues arise, while also being able to gradually shift traffic to the new model. Which deployment strategy should they use?Machine Learning Implementation and Operations
  6. 6.A data engineering team is building a new feature store to manage features for various machine learning models across different projects. They require a solution that can serve features with very low latency (milliseconds) for real-time inference and also provide historical feature values for model training and batch inference. The solution must support point-in-time correctness for reproducible training and backtesting. Which AWS service combination offers these capabilities?Machine Learning Implementation and Operations
  7. 7.A data science team is developing a critical machine learning model for real-time anomaly detection. They need to ensure that the model is always available and highly resilient to instance failures or Availability Zone (AZ) outages. The application requires very low latency for predictions. Which SageMaker endpoint configuration best addresses these requirements?Machine Learning Implementation and Operations
  8. 8.A data science team has developed a new fraud detection model using Amazon SageMaker. The team wants to deploy this model to a production environment where it needs to handle real-time inference requests with low latency and high availability. The model is relatively small and can be loaded into memory quickly. Which deployment option is most suitable for this scenario?Machine Learning Implementation and Operations
  9. 9.A data science team is developing a new credit scoring model. They want to ensure that the model is fair and unbiased across different demographic groups (e.g., age, gender, ethnicity) and that its decisions are explainable to stakeholders and regulatory bodies. Which Amazon SageMaker feature should they integrate into their MLOps workflow to address these requirements?Machine Learning Implementation and Operations
  10. 10.An e-commerce company uses an ML model for real-time product recommendations. The model is deployed on a SageMaker endpoint. To ensure high availability and fault tolerance, they need to ensure that the inference endpoint can withstand an Availability Zone (AZ) outage. Which SageMaker endpoint configuration best addresses this requirement?Machine Learning Implementation and Operations
  11. 11.A data science team is developing a new credit risk assessment model using Amazon SageMaker. They want to automate the entire machine learning workflow, from data preparation and model training to deployment and monitoring, using AWS services. New model versions should be automatically built and deployed whenever new training data becomes available or model code changes. Which AWS service combination provides the most suitable foundation for building such a CI/CD pipeline for ML models?Machine Learning Implementation and Operations
  12. 12.A financial institution uses a machine learning model for real-time fraud detection. The model has been performing well, but due to evolving fraud patterns, its accuracy has recently started to degrade. The data science team needs to implement a system that automatically detects this performance degradation and triggers an alert for model retraining. Which AWS service is best suited to monitor the model's performance metrics and detect concept drift?Machine Learning Implementation and Operations
  13. 13.An e-commerce company uses an ML model for real-time product recommendations. The model processes millions of inference requests per hour, with traffic fluctuating significantly throughout the day. To manage costs and ensure consistent performance, the ML team needs to dynamically adjust the number of inference instances based on the incoming request load. Which Amazon SageMaker feature should they implement?Machine Learning Implementation and Operations
  14. 14.A financial institution is deploying a new fraud detection model to an Amazon SageMaker real-time endpoint. The model is highly sensitive to latency, and any delay in prediction could lead to significant financial losses. The institution wants to ensure that the model remains available and performs optimally even during unexpected spikes in request traffic. Which SageMaker feature should be implemented to address these requirements?Machine Learning Implementation and Operations
  15. 15.A large e-commerce platform uses an ML model for real-time fraud detection. The model processes millions of transactions per hour, and high availability is paramount. Any downtime or performance degradation could lead to significant financial losses. The team needs to deploy the model in a way that provides automatic failover and resilience against Availability Zone (AZ) outages. Which SageMaker deployment option should they choose?Machine Learning Implementation and Operations
  16. 16.A data science team is deploying a new object detection model to an Amazon SageMaker endpoint. During initial testing, they observe that the endpoint instance takes longer than expected to become 'InService' after deployment or scaling events, occasionally leading to failed deployments. The model's container initializes several large dependencies at startup. Which SageMaker endpoint configuration parameter should they adjust to allow more time for the container to fully initialize and pass health checks?Machine Learning Implementation and Operations
  17. 17.A healthcare provider is deploying a diagnostic ML model that processes sensitive patient data. Regulatory compliance (e.g., HIPAA) requires strict data provenance and an immutable audit trail of every inference request, including the input data, model version, and prediction results. This audit trail is essential for post-hoc analysis, debugging, and demonstrating compliance. Which SageMaker feature, combined with appropriate AWS logging, can best meet these requirements?Machine Learning Implementation and Operations
  18. 18.A machine learning engineer needs to deploy a new model that accurately predicts customer churn. Before fully replacing the existing model, the engineer wants to compare the performance of the new model against the current production model using real-time traffic, but only for a small percentage of users, without impacting the majority. Which SageMaker deployment strategy should the engineer use?Machine Learning Implementation and Operations
  19. 19.A data scientist needs to deploy a PyTorch model to an AWS IoT Greengrass device for local inference. The device has limited computational resources and memory. Before deployment, the model needs to be optimized for performance and size on the target hardware. Which SageMaker feature should the data scientist use to achieve this?Machine Learning Implementation and Operations
  20. 20.A data engineering team is building a new feature store to manage features for various machine learning models across their organization. They need a solution that can serve features with very low latency for real-time inference, as well as provide historical feature data for model training and batch inference. Which SageMaker component or mode of operation of the Feature Store meets both of these requirements?Machine Learning Implementation and Operations
  21. 21.A data science team is deploying a new object detection model to an Amazon SageMaker endpoint. This model is computationally intensive, and the team observes that inference requests sometimes time out under high load, even with sufficient instance scaling. They suspect that the default health check timeout might be too short for the model's initialization or a single inference request. Which SageMaker endpoint configuration parameter should they adjust to allow more time for the container to become healthy or process a request?Machine Learning Implementation and Operations
  22. 22.A data science team has developed a new image classification model using a custom TensorFlow container. They need to deploy this model to a SageMaker real-time endpoint for inference. The model artifact is stored in Amazon S3. Which component is primarily responsible for serving the model predictions at the endpoint?Machine Learning Implementation and Operations
  23. 23.A financial institution uses a machine learning model for real-time fraud detection. The model is deployed on Amazon SageMaker endpoints. Due to regulatory compliance requirements and the sensitive nature of the data, all inference requests and responses must be encrypted in transit and at rest. Additionally, the solution must minimize operational overhead. Which SageMaker feature should be used to ensure encryption of inference data in transit?Machine Learning Implementation and Operations
  24. 24.An e-commerce company uses a fraud detection model that needs to process millions of transactions daily. The model's predictions are critical, and any downtime or high latency could result in significant financial losses. The company requires a highly available and fault-tolerant deployment strategy for its SageMaker endpoint, ensuring continuous operation even during instance failures or updates. Which SageMaker endpoint configuration best meets these requirements?Machine Learning Implementation and Operations
  25. 25.A data science team has developed a new image classification model using a custom TensorFlow version and specific pre-processing libraries not available in standard SageMaker images. They need to deploy this model to a SageMaker real-time endpoint for inference. How can they package their model and dependencies to ensure it runs correctly on the SageMaker endpoint?Machine Learning Implementation and Operations
  26. 26.A data science team is building a personalized content recommendation system. They need to frequently update their models (daily retrains) and deploy new versions automatically to production after successful testing. The entire process, from data ingestion to model deployment, must be automated, repeatable, and trackable. They also want to include steps for data preprocessing, model training, evaluation, and conditional deployment based on evaluation metrics. Which AWS service provides the orchestration capabilities for such an end-to-end MLOps workflow?Machine Learning Implementation and Operations
  27. 27.A data engineering team is setting up a Feature Store to manage features for various machine learning models. They need to ensure that the features are consistently available for both online (low-latency inference) and offline (batch training) use cases. Which aspect of a Feature Store architecture directly addresses this dual requirement?Machine Learning Implementation and Operations
  28. 28.A pharmaceutical company is developing a machine learning model to assist in drug discovery, which involves extremely large and complex deep learning models. The inference latency is critical, but the cost of deploying full GPU instances for every inference request is prohibitive. The model needs to run on SageMaker, and the team wants to optimize cost-performance for inference by leveraging GPU acceleration only when necessary, attached to CPU instances. Which SageMaker feature should they utilize?Machine Learning Implementation and Operations
  29. 29.A large e-commerce platform uses an ML model for real-time product recommendations. The model processes millions of inference requests per hour, with traffic patterns that fluctuate significantly throughout the day. To ensure responsiveness and optimize costs, the platform needs to automatically scale the SageMaker endpoint instances based on the incoming traffic load. Which SageMaker feature should be configured?Machine Learning Implementation and Operations
  30. 30.A data science team is developing a critical machine learning model for real-time anomaly detection. They are using an AWS CodePipeline-based MLOps pipeline for continuous integration and continuous delivery. A new model version is ready for deployment. The team wants to ensure that the new model undergoes a rigorous performance and stability check in a production-like environment, processing actual production traffic, but without directly impacting the real-time responses to end-users. Only after a successful observation period will the model be considered for full deployment. Which deployment strategy should be integrated into the CodePipeline?Machine Learning Implementation and Operations
  31. 31.A data science team is developing a fraud detection model. They need a centralized, version-controlled repository to store, share, and manage features used across multiple models and teams, ensuring consistency between training and inference. The solution must support both online (low-latency) and offline (batch) access patterns. Which AWS service is best suited for this requirement?Machine Learning Implementation and Operations
  32. 32.A startup is building a personalized content recommendation system. They need to frequently update the recommendations based on user interactions and content changes, requiring daily model retraining. The entire MLOps workflow, from data ingestion and processing to model training, evaluation, and deployment, needs to be fully automated and orchestrated. Which AWS service is best suited for defining and orchestrating this end-to-end machine learning workflow?Machine Learning Implementation and Operations
  33. 33.A company is using a deep learning model for image classification deployed on an Amazon SageMaker endpoint. To reduce inference costs and improve throughput, they want to optimize the model for specific hardware accelerators available on SageMaker instances. The optimization process needs to be framework-agnostic and produce a deployable artifact. Which SageMaker capability should they leverage?Machine Learning Implementation and Operations
  34. 34.A pharmaceutical company is developing a machine learning model to assist in drug discovery. The model is computationally intensive, requiring significant GPU resources for inference, but the inference requests are infrequent and latency is not extremely critical (a few seconds is acceptable). The company wants to optimize costs for these inference workloads. Which SageMaker feature should they consider to achieve cost-effective GPU-powered inference?Machine Learning Implementation and Operations
  35. 35.A media company uses a content moderation model deployed on a SageMaker real-time endpoint. They want to introduce a new, improved version of the model to a small, controlled group of users (e.g., 5%) to observe its real-world performance and stability before a full rollout. They need to be able to quickly revert to the old model if any major issues arise, minimizing impact on the majority of users. Which deployment strategy should they employ?Machine Learning Implementation and Operations
  36. 36.A large e-commerce platform uses a recommendation engine powered by a machine learning model deployed on Amazon SageMaker. During peak shopping seasons, the inference requests can surge unexpectedly, leading to increased latency and potential service disruptions. The platform needs a solution that can automatically adjust the endpoint's capacity to handle these unpredictable traffic spikes without manual intervention and ensure consistent low latency. Which combination of SageMaker features should be implemented?Machine Learning Implementation and Operations
  37. 37.An e-commerce company uses an ML model for real-time product recommendations. The model processes millions of inference requests per hour. Due to the high volume and the need for auditing and debugging, the company requires that all inference requests and responses, including the input and output data, be captured and stored for future analysis. Which SageMaker feature enables this capability?Machine Learning Implementation and Operations
  38. 38.A financial institution is developing a new credit risk assessment model. They need a centralized, versioned, and low-latency storage solution for features that can be used consistently across both model training and real-time inference. The solution must support point-in-time correctness for historical training data and fast retrieval for online predictions. Which SageMaker service is designed to meet these requirements?Machine Learning Implementation and Operations
  39. 39.A machine learning engineer is deploying a real-time inference endpoint for a critical fraud detection model using Amazon SageMaker. The model requires high availability and needs to be resilient to instance failures within a single Availability Zone. Which configuration should the engineer primarily focus on to address this requirement?Machine Learning Implementation and Operations
  40. 40.A retail company is deploying a new personalized recommendation model. To minimize risk, they want to gradually shift traffic from the old model (Model A) to the new model (Model B), carefully monitoring Model B's performance and impact on key business metrics like conversion rate. If Model B performs worse than expected, they need to be able to quickly revert to Model A. Which deployment strategy should they employ?Machine Learning Implementation and Operations
  41. 41.A data science team is building a large language model (LLM) for a natural language processing application. The model is too large to fit into the memory of a single GPU instance or to process efficiently on a single instance during inference. The team needs to deploy this LLM on Amazon SageMaker, optimizing for memory utilization and inference latency across multiple instances. Which SageMaker feature is specifically designed to handle the deployment of very large models by distributing them across multiple devices or instances?Machine Learning Implementation and Operations
  42. 42.A data science team is developing a new credit risk assessment model. They need a centralized, versioned repository to store, manage, and share features across different machine learning models and teams. This repository must support both online inference and offline training. Which AWS service is best suited for this requirement?Machine Learning Implementation and Operations
  43. 43.A manufacturing company uses a machine learning model for predictive maintenance on factory equipment. The model is deployed on SageMaker. Over time, the types of equipment failures and their underlying causes evolve due to new machinery, production processes, and environmental factors. The current model, trained on historical data, fails to accurately predict these new failure modes, leading to increased downtime. This phenomenon is known as concept drift. What is the most effective MLOps strategy to address this specific challenge?Machine Learning Implementation and Operations
  44. 44.A data science team needs to implement a robust MLOps practice where new model versions are automatically trained, evaluated, and deployed based on new data or code changes. They want to define a series of steps for this entire workflow, including data preprocessing, training, model evaluation, and conditional deployment, all orchestrated within SageMaker. Which AWS service is purpose-built for defining and managing such end-to-end ML workflows?Machine Learning Implementation and Operations
  45. 45.A financial services company is developing a new credit risk assessment model. They need to ensure that the model's predictions are fair and unbiased across different demographic groups (e.g., age, gender) and that the model's decision-making process is transparent and explainable to comply with regulatory requirements. Which Amazon SageMaker feature should the data science team use to achieve these goals during model development and post-deployment monitoring?Machine Learning Implementation and Operations
  46. 46.A data science team is developing a new machine learning model for a critical real-time application. They need to implement a robust MLOps practice where new model versions are automatically built, tested, and deployed to production, ensuring consistency and reproducibility across the entire ML workflow. The process should include steps for data preprocessing, model training, model evaluation, and conditional deployment based on performance metrics. Which SageMaker feature provides the most comprehensive solution for orchestrating this end-to-end workflow?Machine Learning Implementation and Operations
  47. 47.A media company uses a machine learning model to personalize content recommendations for its users. The model is deployed on an Amazon SageMaker endpoint. Over time, the performance of the model has degraded, and the data science team suspects that the underlying data distribution has changed, leading to 'concept drift'. To systematically address this, they want to automatically retrain the model when significant data changes are detected. Which SageMaker service can be used to monitor data quality and trigger model retraining based on drift detection?Machine Learning Implementation and Operations
  48. 48.A manufacturing company uses a machine learning model to predict equipment failures. The model was trained using historical sensor data. They want to deploy this model to edge devices on the factory floor, where internet connectivity is intermittent and processing power is limited. They need to optimize the model for size and performance on these resource-constrained devices. Which SageMaker capability is best suited for this specific optimization and deployment scenario?Machine Learning Implementation and Operations
  49. 49.A machine learning engineer is deploying a large language model (LLM) on Amazon SageMaker. The model is too large to fit into the memory of a single GPU instance for inference. The engineer needs to optimize inference performance while ensuring the model can be served efficiently. Which SageMaker feature should the engineer use to distribute the model across multiple GPUs or instances for inference?Machine Learning Implementation and Operations
  50. 50.A pharmaceutical company is developing a machine learning model to assist in drug discovery. The model is computationally intensive and requires significant GPU resources for inference. To optimize costs while maintaining an acceptable latency of less than 100ms, they want to use a shared, elastic GPU resource for inference that can be attached to their CPU-based SageMaker instances. Which SageMaker feature should they use?Machine Learning Implementation and Operations