AWS Certified Machine Learning – Specialty practice questions

243 free questions with answers and explanations.

Practice test
  1. 101.A startup is building a personalized content recommendation system. They need to frequently experiment with new model architectures and feature sets, deploy them quickly, and gather immediate feedback on their performance against existing models in a production environment. Which MLOps practice would best facilitate this iterative experimentation and evaluation?Machine Learning Implementation and Operations
  2. 102.A financial institution is deploying a new credit risk assessment model. Before fully rolling it out to all customers, they want to compare its performance against the existing legacy model in a production environment using a small subset of live traffic. They need to ensure that the new model's predictions are captured for analysis without impacting the customer experience of the existing model. Which A/B testing strategy should they use?Machine Learning Implementation and Operations
  3. 103.A machine learning engineer is tasked with optimizing the inference performance of a large, complex model deployed on a SageMaker endpoint. The model has many layers, and the bottleneck is identified as the sequential execution of these layers. The engineer wants to parallelize parts of the inference process to reduce latency. Which advanced inference optimization technique is most appropriate for this scenario?Machine Learning Implementation and Operations
  4. 104.A research institution is collecting sensor data from remote environmental monitoring stations. The data is small in size (a few kilobytes per reading) but arrives at a very high frequency (thousands of readings per second) from thousands of devices. The institution needs to ingest this data efficiently, store it for long-term analysis, and prepare it for machine learning models that predict environmental changes. Which AWS service is most cost-effective and scalable for collecting and ingesting this high-volume, low-latency data stream?Data Engineering
  5. 105.A data engineer needs to ensure that sensitive customer data, stored in an Amazon S3 data lake, is only accessible by authorized machine learning models and specific data scientists. The data is encrypted at rest using KMS. To enforce fine-grained access control based on user roles and specific S3 prefixes (folders), which AWS service should be primarily used?Data Engineering
  6. 106.A data engineer is working with a large dataset stored in Amazon S3, consisting of billions of records. For efficient querying and to improve the performance of downstream machine learning model training, the data needs to be partitioned. The most frequent queries involve filtering data by `event_date` and `customer_region`. Which partitioning strategy should the data engineer implement to optimize query performance and reduce data scanning costs?Data Engineering
  7. 107.A financial institution is analyzing credit card transaction data to detect fraudulent activities. They are particularly interested in understanding the distribution of transaction amounts. The data shows a heavily right-skewed distribution with a long tail of very large transactions. Which statistical measure of central tendency would be least affected by these extreme values and best represent the 'typical' transaction amount?Exploratory Data Analysis
  8. 108.A data scientist is working with a large dataset of customer reviews, which contains free-form text. Before feeding this data into a natural language processing (NLP) model, the text needs to be cleaned by removing special characters, converting to lowercase, tokenizing, and removing common stopwords. This preprocessing step must be easily repeatable and scalable. Which AWS service is best suited for performing these common text data preparation tasks in a managed and scalable way?Data Engineering
  9. 109.A data scientist is preparing a large dataset for training a machine learning model. The dataset contains several categorical features with high cardinality, some of which exhibit a skewed distribution where a few categories account for a vast majority of the observations, while many others are rare. The data scientist needs to transform these features to improve model performance and reduce the dimensionality without losing significant information. Which feature engineering technique is most appropriate for this scenario?Data Engineering
  10. 110.A research institution collects genomic sequence data, which is extremely large (petabytes) and highly sensitive. This data is stored in Amazon S3. Due to regulatory compliance, the data must be encrypted at rest and in transit, and access must be strictly controlled with granular permissions, ensuring only authorized researchers can access specific subsets of the data. Furthermore, all access attempts, successful or not, must be logged for auditing purposes. Which combination of AWS security features provides the most comprehensive solution for this scenario?Data Engineering
  11. 111.A machine learning engineer is tasked with building a model to predict equipment failure based on sensor readings. During EDA, they notice that several sensor readings ('temperature', 'vibration', 'pressure') are highly correlated with each other, exhibiting multicollinearity. This can lead to unstable model coefficients and reduced interpretability. What is the most effective data cleaning/feature engineering technique to address this multicollinearity while retaining as much information as possible from the original features?Exploratory Data Analysis
  12. 112.A global e-commerce company wants to analyze customer behavior across different regions and languages. Their customer data, stored in Amazon S3, is currently organized by `upload_date`. To facilitate efficient queries and machine learning model training that frequently filter data by `country` and `language`, the data engineering team decides to re-organize the data. Which data partitioning scheme in S3 would be most effective for this scenario?Data Engineering
  13. 113.A data scientist is analyzing customer churn data. They want to visualize the relationship between 'MonthlyCharges' (continuous) and 'TotalCharges' (continuous) for churned vs. non-churned customers. They suspect there might be a non-linear relationship or varying density patterns within these groups. Which data visualization technique would best reveal these insights?Exploratory Data Analysis
  14. 114.A machine learning team is developing a credit risk model. They have a large dataset containing customer transaction history, and they need to create a feature representing the 'total transaction amount in the last 30 days' for each customer. This feature needs to be updated daily. Which feature engineering technique is described here, and what AWS service is suitable for its efficient computation on a large dataset?Data Engineering
  15. 115.A data engineer is integrating a new data source into an existing data lake. The new source provides daily CSV files, each containing hundreds of millions of records. These files need to be converted to Parquet format, partitioned by 'ingestion_date' and 'source_id', and then registered in the AWS Glue Data Catalog. The process must be scheduled to run daily. Which AWS service provides a serverless and managed way to orchestrate this entire workflow, including data conversion, partitioning, and catalog updates?Data Engineering
  16. 116.A data scientist is training a fraud detection model using a large, imbalanced dataset where fraud cases are rare. To improve model performance, they decide to use Synthetic Minority Oversampling Technique (SMOTE) to balance the classes. This process needs to run on a large dataset and integrate into an existing SageMaker pipeline. Which AWS service is best suited for executing this computationally intensive data preparation step within the SageMaker ecosystem?Data Engineering
  17. 117.A data engineering team is integrating data from various sources into a unified data lake. They encounter a 'CustomerID' field that is represented as an integer in one source (e.g., 12345) and as a string with leading zeros in another (e.g., '0012345'). Both represent the same unique identifier. To ensure data consistency and enable proper joins and analysis, which data cleaning operation should be applied?Exploratory Data Analysis
  18. 118.A data scientist is preparing a dataset for a natural language processing (NLP) model. The dataset consists of millions of customer reviews, and a significant portion contains personally identifiable information (PII) such as names, email addresses, and phone numbers. Before training the model, this PII must be removed or masked to comply with privacy regulations. Which AWS service is explicitly designed to detect and redact sensitive data from text at scale?Data Engineering
  19. 119.A financial institution needs to analyze customer transaction data to detect fraudulent activities. The data arrives from various online and offline channels in near real-time. The volume can spike significantly during peak hours. The solution requires processing these incoming transactions, enriching them with customer master data from a database, and then making the enriched data available for real-time fraud detection models. The solution must be serverless, highly scalable, and capable of handling varying throughput. Which AWS service combination would best meet these requirements?Data Engineering
  20. 120.A data scientist is analyzing a dataset containing customer ratings for various products on an e-commerce platform. The ratings are on a scale of 1 to 5. They want to understand the central tendency and spread of these ratings. However, they observe that the distribution of ratings is heavily concentrated at 5 stars, with a long tail of lower ratings (left-skewed). They need a robust measure of central tendency and spread that is not unduly influenced by this skewness. Which pair of statistical measures would be most appropriate?Exploratory Data Analysis
  21. 121.A data scientist is performing a rigorous hypothesis test on a new advertising campaign's effectiveness. They have a control group and an experimental group, and the outcome variable ('ConversionRate') is continuous and approximately normally distributed. After running the test, they obtain a p-value of 0.03. The chosen significance level (alpha) for this study is 0.05. Which of the following is the correct conclusion?Exploratory Data Analysis
  22. 122.A research institution is collecting highly sensitive genomic data from various sources. They need a secure and auditable method to ingest this data into an S3 data lake, ensuring that all data transfers are encrypted end-to-end and access is logged. The data sources are often on-premises and generate large files. Which AWS service and configuration combination provides the most secure and auditable data collection for this scenario?Data Engineering
  23. 123.A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a 'customer_age' column, which has a significant number of missing values. The data scientist observes that the distribution of existing 'customer_age' values is heavily skewed towards younger customers (ages 20-35) with a long tail extending to older ages. Which imputation strategy is most appropriate to handle the missing 'customer_age' values while preserving the original distribution characteristics?Data Engineering
  24. 124.A data scientist is analyzing a dataset of customer demographics. They want to check if there is a statistically significant association between 'PreferredCommunicationChannel' (categorical: Email, SMS, Phone) and 'PurchaseFrequencyGroup' (categorical: Low, Medium, High). Which statistical test is most appropriate for this analysis?Exploratory Data Analysis
  25. 125.A data scientist is working with a time-series dataset where observations are recorded hourly. Missing values are present in key sensor readings. For a predictive model, it's crucial to fill these missing values in a way that reflects the typical daily and weekly patterns, rather than just simple linear interpolation or forward-fill. The dataset spans several years, and the missing data can occur in short bursts or longer periods. Which imputation strategy is most appropriate for capturing these cyclical patterns?Data Engineering
  26. 126.A data engineering team is building a new machine learning pipeline that processes sensitive customer data. They need to ensure that the data is encrypted both in transit and at rest within Amazon S3 to meet compliance requirements. Additionally, they want to manage their own encryption keys to maintain full control over the data. Which encryption method should the team choose for their Amazon S3 buckets?Data Engineering
  27. 127.A data scientist is analyzing a large dataset of customer feedback text. They want to identify groups of reviews that express similar themes or topics without any pre-labeled data. The goal is to uncover latent topics within the text to better understand customer concerns. Which unsupervised technique is most suitable for this task?Exploratory Data Analysis
  28. 128.A data scientist is analyzing a large dataset of customer behavior, including features like 'browsing_duration', 'number_of_clicks', and 'purchase_amount'. During data cleaning, they identify a small number of records where 'browsing_duration' is recorded as 0, but 'number_of_clicks' is very high (e.g., 500+). These records contradict logical expectations (zero duration implies no activity, thus zero clicks). Removing these records would lead to significant data loss for other valuable features. Which data cleaning strategy is most appropriate for these contradictory records?Exploratory Data Analysis
  29. 129.A data scientist needs to prepare a dataset for a time-series forecasting model. The dataset contains sensor readings collected at irregular intervals, resulting in missing values. To ensure the model can be trained effectively, the data scientist needs to fill these missing values using a method that considers the temporal order and trends in the data. Which data preparation technique is most appropriate for handling missing values in this time-series context?Data Engineering
  30. 130.A data engineer is working with a large transactional dataset. They observe that some transactions have extremely high values compared to the vast majority, potentially indicating fraudulent activity or data entry errors. These extreme values are several standard deviations away from the mean and are significantly distorting statistical measures like the mean and standard deviation. What is the most appropriate statistical method to identify these potential outliers?Exploratory Data Analysis
  31. 131.A data scientist is preparing a dataset for an image classification model. The dataset contains millions of images, each with varying resolutions and aspect ratios. To ensure consistent input for the neural network and to prevent distortion, the images need to be resized and then cropped to a fixed square dimension (e.g., 224x224 pixels) while maintaining the aspect ratio as much as possible before cropping. Which sequence of image transformation operations should be applied?Data Engineering
  32. 132.A machine learning engineer is analyzing sensor data from an industrial machine. The 'Temperature' readings are consistently positive but exhibit a highly skewed distribution with a long tail towards higher temperatures, indicating occasional overheating events. The engineer wants to transform this data to achieve a more symmetric, Gaussian-like distribution to meet the assumptions of certain statistical models. Which transformation is most appropriate for this type of data?Exploratory Data Analysis
  33. 133.A data science team is analyzing customer feedback text data to understand sentiment. They have collected over 100,000 comments. Before building a sentiment analysis model, they need to perform Exploratory Data Analysis (EDA). Which visualization technique would be most effective for quickly identifying the most frequently occurring positive and negative keywords in the entire corpus?Exploratory Data Analysis
  34. 134.A data engineering team needs to transform raw streaming data from IoT devices into a structured format suitable for machine learning. The data arrives as JSON objects, but some fields are nested, and inconsistent schema versions occasionally appear. They need to flatten the nested structures, handle schema evolution, and convert the data into Apache Parquet format before storing it in Amazon S3. Which AWS service is best suited for performing these transformations in a serverless, scalable, and cost-effective manner?Data Engineering
  35. 135.A financial company needs to store petabytes of historical transaction data for compliance and future machine learning model training. The data is accessed infrequently, but when accessed, queries often involve scanning large portions of the dataset. Cost-efficiency and durability are paramount, with eventual consistency being acceptable. Which Amazon S3 storage class is most appropriate for this requirement?Data Engineering
  36. 136.A data engineer is working with a large dataset of customer reviews stored in Amazon S3. The reviews contain various emojis, special characters, and HTML tags that need to be removed before the data can be used for sentiment analysis. The dataset is several terabytes in size and needs to be processed efficiently. Which AWS service is most suitable for performing this large-scale text cleaning and preprocessing?Data Engineering
  37. 137.A data scientist is analyzing a dataset of customer survey responses. One question asks for 'CustomerSatisfaction' on a scale of 1 to 5. They notice that the responses are heavily concentrated at 4 and 5, with very few responses at 1, 2, or 3. They want to check if this observed distribution significantly differs from a hypothetical uniform distribution (where each rating 1-5 has an equal chance). Which statistical test is most appropriate for this comparison?Exploratory Data Analysis
  38. 138.A data engineering team is building a data lake on Amazon S3 for an analytics platform. They receive data from various sources in different formats (CSV, JSON, Parquet). To enable efficient querying and machine learning model training without rewriting queries for each format, they need a centralized metadata repository that describes the schema, location, and partition information of all datasets in the data lake. This repository should also allow integration with services like Amazon Athena and AWS Glue. Which AWS service is best suited for this purpose?Data Engineering
  39. 139.A data engineering team is building a machine learning pipeline that processes highly sensitive customer data. Before training models, PII (Personally Identifiable Information) must be masked or anonymized in the dataset. The data is stored in Amazon S3 as CSV files and needs to be processed in a serverless, scalable manner. The masking process involves replacing specific columns (e.g., 'customer_name', 'email_address') with fictitious but consistent values (e.g., 'customer_name_1', 'customer_name_2') across different rows for the same original PII, to maintain referential integrity for downstream analysis. Which approach is most suitable?Data Engineering
  40. 140.A data scientist is analyzing sensor data from industrial machinery. The data contains several time-series features, and they suspect that some sensors occasionally report erroneous, extremely high values that are physically impossible and occur as sudden, isolated spikes. These spikes are not part of any normal operational variance and can severely distort subsequent analysis. Which advanced outlier detection technique is best suited to identify these specific types of outliers in time-series data?Exploratory Data Analysis
  41. 141.A data scientist is analyzing a dataset of customer demographics. They want to visualize the distribution of 'age' and 'income' simultaneously to identify any potential clusters or relationships. The 'age' feature is normally distributed, while 'income' is heavily skewed to the right. Which data visualization technique is most appropriate to display the joint distribution of these two features effectively, considering their different distributional properties?Exploratory Data Analysis
  42. 142.A data scientist is performing Exploratory Data Analysis (EDA) on a dataset of customer demographics for a new product launch. They have collected data on 'age', 'gender', 'income_level' (low, medium, high), and 'purchase_intent' (binary: yes/no). The team wants to understand if there is a statistically significant association between 'income_level' and 'purchase_intent'. Which hypothesis test is the most appropriate to determine this association?Exploratory Data Analysis
  43. 143.A data engineer is designing a data ingestion pipeline for real-time sensor data from thousands of IoT devices. Each device sends small JSON payloads every few seconds. The data needs to be processed, transformed, and then stored in a data lake for analytics and machine learning. The solution must be highly scalable, serverless, and cost-effective. Which combination of AWS services should be used for this pipeline?Data Engineering
  44. 144.A data scientist is preparing a dataset for a machine learning model. They notice that the 'CustomerAge' column contains several negative values, which are logically impossible. Which of the following data cleaning techniques is most appropriate to address this issue?Exploratory Data Analysis
  45. 145.A data engineering team is designing a new machine learning pipeline that will process sensitive customer data. They need to ensure that the data remains encrypted both at rest and in transit, and that access is strictly controlled. Which combination of AWS services and features provides the most robust solution for securing this data within an Amazon S3 data lake?Data Engineering
  46. 146.A startup is building a recommendation engine and needs to process petabytes of clickstream data from its website. This data arrives continuously, and transformations, such as sessionization and feature extraction, must be applied in near real-time before being stored for model training. The solution must be serverless, highly scalable, and cost-effective. Which AWS service combination should the startup use to achieve this?Data Engineering
  47. 147.A data scientist is analyzing a large dataset of customer reviews. They want to identify groups of customers with similar sentiment patterns without pre-defining the number of groups. The sentiment scores are continuous values between -1 (negative) and 1 (positive). Which unsupervised machine learning technique is best suited for this task?Exploratory Data Analysis
  48. 148.A data engineer is tasked with optimizing the query performance for a large dataset of several petabytes stored in Amazon S3, which is used by Amazon Athena for ad-hoc analytics. The data is currently stored as uncompressed JSON files, and queries often involve filtering by specific date ranges and customer segments. The current query times are unacceptably long and expensive. Which combination of data partitioning and file format conversion would significantly improve query performance and reduce costs?Data Engineering
  49. 149.A healthcare organization is setting up a new data pipeline for machine learning, handling sensitive patient data. Due to HIPAA compliance requirements, all data must be pseudonymized before it is used for model training. This involves replacing direct identifiers (e.g., patient names, social security numbers) with artificial identifiers while maintaining referential integrity for analytical purposes. Which data transformation technique is most appropriate to meet this compliance requirement?Data Engineering
  50. 150.A data engineering team is building a feature store for a machine learning pipeline. They have ingested a massive dataset, and during the initial Exploratory Data Analysis (EDA), they discover that a critical numerical feature, 'customer_lifetime_value', has a long tail of extremely high values (positive skew) and also contains a significant number of zero values, representing customers with no recorded value yet. A standard logarithmic transformation (log(x)) would fail for zero values. To normalize this feature for model training while preserving its relative relationships and handling zeros, which transformation strategy is most appropriate?Exploratory Data Analysis