Data Engineering
Preparing and managing data for machine learning.
Getting Started: Exam Overview
Free knowledge base
Everything from the course in one searchable place: 211 entries. Use it to review before a practice test or look up a word you forgot.
211 results
Preparing and managing data for machine learning.
Getting Started: Exam Overview
Analyzing datasets to summarize their main characteristics.
Getting Started: Exam Overview
Selecting, training, and evaluating machine learning algorithms.
Getting Started: Exam Overview
Practices for deploying and maintaining ML models in production.
Getting Started: Exam Overview
Optimizing model configuration parameters for best performance.
Getting Started: Exam Overview
Creating new input features from existing data to improve models.
Getting Started: Exam Overview
When model performance degrades over time due to data changes.
Getting Started: Exam Overview
To remember the exam domains: D.E.M.O. - Data Engineering, Exploratory Data Analysis, Modeling, Operations (ML Implementation and Operations).
Getting Started: Exam Overview
On the exam, pay close attention to scenario-based questions that ask you to choose the 'MOST appropriate' or 'BEST' service or solution for a given ML problem. Keywords like 'real-time inference,' 'batch processing,' 'large datasets,' or 'cost-effective' are critical hints.
Getting Started: Exam Overview
Underestimating the importance of Data Engineering and EDA – these foundational steps are crucial for model success.
Getting Started: Exam Overview
Focusing solely on theoretical ML concepts without understanding their practical application on AWS services.
Getting Started: Exam Overview
Neglecting MLOps practices, which are critical for real-world model deployment and maintenance.
Getting Started: Exam Overview
Question type with one correct answer among several options.
Getting Started: Exam Overview
Question type requiring selection of all correct answers.
Getting Started: Exam Overview
Method converting raw scores to a standardized scale (100-1000).
Getting Started: Exam Overview
Official AWS document outlining exam domains and objectives.
Getting Started: Exam Overview
Incorrect options in a multiple-choice question.
Getting Started: Exam Overview
Questions describing real-world problems to test application of knowledge.
Getting Started: Exam Overview
To remember the passing score: 'Seven Fifty' is the key to your ML 'Specialty' success!
Getting Started: Exam Overview
The exam requires a scaled score of 750 or higher to pass. Memorize that there are 65 questions and 170 minutes allowed for the exam.
Getting Started: Exam Overview
Not reading the entire question and all answer options before selecting an answer.
Getting Started: Exam Overview
Underestimating the importance of hands-on experience with AWS ML services.
Getting Started: Exam Overview
Failing to manage time effectively during the exam, spending too long on difficult questions.
Getting Started: Exam Overview
A centralized repository for raw, unstructured data.
Data Engineering Fundamentals
Continuous flow of data generated in real-time.
Data Engineering Fundamentals
Extract, Transform, Load; process to move and prepare data.
Data Engineering Fundamentals
Scaling numerical features to a standard range.
Data Engineering Fundamentals
Filling in missing data values, often with estimates.
Data Engineering Fundamentals
Scalable object storage for any type of data.
Data Engineering Fundamentals
Serverless data integration service for ETL.
Data Engineering Fundamentals
CST: Collect, Store, Transform. Remember the order of operations for data preparation!
Data Engineering Fundamentals
The exam frequently tests your knowledge of which AWS service to use for specific data collection, storage, or transformation tasks. Keywords like 'real-time streaming,' 'large-scale batch ETL,' 'unstructured data lake,' or 'relational database' should immediately point you to services like Kinesis, Glue, S3, or RDS respectively.
Data Engineering Fundamentals
Underestimating the complexity and time required for data cleaning and transformation.
Data Engineering Fundamentals
Choosing a storage solution that doesn't scale with data growth or meet access pattern needs.
Data Engineering Fundamentals
Neglecting data governance and security during collection and storage phases.
Data Engineering Fundamentals
Converting categorical data into a binary vector representation.
Data Engineering Fundamentals
Assigning a unique integer to each category in a categorical feature.
Data Engineering Fundamentals
Scaling data to have a mean of 0 and a standard deviation of 1 (Z-score).
Data Engineering Fundamentals
Scaling data to a fixed range, typically between 0 and 1.
Data Engineering Fundamentals
New features created by combining two or more existing features.
Data Engineering Fundamentals
Grouping continuous numerical values into discrete intervals or bins.
Data Engineering Fundamentals
FEAT: Features Engineered Are Terrific! It's all about making your data look its best for the model.
Data Engineering Fundamentals
The exam frequently tests your understanding of when to use specific encoding or scaling techniques. Remember to standardize data for algorithms sensitive to feature magnitudes (e.g., SVMs, K-Means, Neural Networks) and use one-hot encoding for nominal categorical features to avoid implying an ordinal relationship.
Data Engineering Fundamentals
Applying scaling techniques to target variables instead of only features.
Data Engineering Fundamentals
Using label encoding for nominal categorical features, which can imply an incorrect ordinal relationship.
Data Engineering Fundamentals
Not handling missing values before training, leading to errors or poor model performance.
Data Engineering Fundamentals
Dividing data into smaller parts based on column values for faster queries.
Data Engineering Fundamentals
A persistent metadata store for data assets across AWS services.
Data Engineering Fundamentals
Automatically discovers data, infers schemas, and populates the Data Catalog.
Data Engineering Fundamentals
The process of adapting to changes in data structure over time.
Data Engineering Fundamentals
Serverless query service that uses SQL to analyze data in S3 via Glue Catalog.
Data Engineering Fundamentals
Redshift feature to query data directly from S3 using Glue Catalog.
Data Engineering Fundamentals
The column(s) used to organize and divide data into partitions.
Data Engineering Fundamentals
P-A-R-T-Y: Partitioning for Analytics, Reductions in cost, Timely queries, Your data cataloged.
Data Engineering Fundamentals
The AWS Certified Machine Learning – Specialty exam frequently tests your understanding of how AWS Glue Data Catalog integrates with other services like Athena, Redshift Spectrum, and EMR. Be prepared to identify scenarios where Glue Crawlers are the appropriate tool for schema discovery and management, especially with S3 data lakes. Remember that Glue Data Catalog is a *managed* service.
Data Engineering Fundamentals
Not partitioning data at all, leading to full table scans and high costs.
Data Engineering Fundamentals
Choosing partition keys that are not frequently used in queries, negating performance benefits.
Data Engineering Fundamentals
Ignoring schema evolution, causing downstream applications to break when data structure changes.
Data Engineering Fundamentals
AWS secures the cloud, you secure in the cloud.
Data Engineering Fundamentals
Managed service for creating and controlling encryption keys.
Data Engineering Fundamentals
Manages access to AWS services and resources securely.
Data Engineering Fundamentals
Granting only the necessary permissions for a task.
Data Engineering Fundamentals
Private connection to AWS services from your VPC.
Data Engineering Fundamentals
Logs API calls and events for auditing and compliance.
Data Engineering Fundamentals
Discovers and protects sensitive data in S3 buckets.
Data Engineering Fundamentals
Data stored in persistent storage, like S3 or databases.
Data Engineering Fundamentals
KMS for Keys, IAM for Access, CloudTrail for Clues, Macie for Mysteries!
Data Engineering Fundamentals
The exam frequently tests the Shared Responsibility Model. Remember: AWS is 'security *of* the cloud,' and the customer is 'security *in* the cloud.' Also, know that KMS is the primary service for encryption key management.
Data Engineering Fundamentals
Forgetting to encrypt data at rest, leaving S3 buckets unencrypted.
Data Engineering Fundamentals
Granting overly permissive IAM roles (e.g., S3:* access) instead of applying the principle of least privilege.
Data Engineering Fundamentals
Not enabling CloudTrail or CloudWatch logging, making it impossible to audit security events.
Data Engineering Fundamentals
Graphical representation of data to reveal patterns and insights.
Exploratory Data Analysis Techniques
Process of detecting and correcting inaccurate or corrupt records in a dataset.
Exploratory Data Analysis Techniques
Absence of data for a particular observation or variable.
Exploratory Data Analysis Techniques
Data point significantly different from other observations.
Exploratory Data Analysis Techniques
Identical entries in a dataset that represent the same entity.
Exploratory Data Analysis Techniques
VCR: Visualize, Clean, Repeat! It's like watching a movie of your data, pausing to fix glitches, then playing again to see the improved picture.
Exploratory Data Analysis Techniques
The exam often tests your ability to choose the right visualization for a given data type or problem (e.g., 'Which plot for distribution?' or 'Which for correlation?'). Also, be ready for scenarios describing data quality issues and asking for the best cleaning strategy.
Exploratory Data Analysis Techniques
Ignoring data visualization and jumping straight to modeling, missing critical data issues.
Exploratory Data Analysis Techniques
Applying a single cleaning technique (e.g., always dropping rows with missing values) without considering the data's context.
Exploratory Data Analysis Techniques
Not documenting cleaning steps, making the process irreproducible and hard to audit.
Exploratory Data Analysis Techniques
Missingness unrelated to any data, observed or unobserved.
Exploratory Data Analysis Techniques
Missingness depends on observed data, not the missing data itself.
Exploratory Data Analysis Techniques
Missingness depends on the value of the missing data itself.
Exploratory Data Analysis Techniques
Measures how many standard deviations a data point is from the mean.
Exploratory Data Analysis Techniques
Range between the first and third quartiles, used for outlier detection.
Exploratory Data Analysis Techniques
Capping extreme values at a specified percentile.
Exploratory Data Analysis Techniques
To remember the types of missing data: 'M&M's Are Not Random' - MCAR, MAR, MNAR. Each 'M' reminds you of the increasing complexity of missingness.
Exploratory Data Analysis Techniques
The exam often tests your understanding of when to use specific imputation techniques (mean/median for numerical, mode for categorical, advanced for complex patterns) and the impact of outliers on different model types (linear models are sensitive, tree-based models are more robust). Be prepared to differentiate between MCAR, MAR, and MNAR.
Exploratory Data Analysis Techniques
Blindly dropping rows with missing values without understanding the missingness mechanism, leading to data loss and bias.
Exploratory Data Analysis Techniques
Using simple mean/median imputation for all scenarios, which can reduce variance and distort relationships, especially for MAR or MNAR data.
Exploratory Data Analysis Techniques
Automatically removing all detected outliers without investigating their domain context, potentially discarding valuable information or legitimate extreme events.
Exploratory Data Analysis Techniques
Summarize and describe features of data.
Exploratory Data Analysis Techniques
Make predictions about a population from a sample.
Exploratory Data Analysis Techniques
States no effect or no difference.
Exploratory Data Analysis Techniques
States an effect or a difference exists.
Exploratory Data Analysis Techniques
Probability of observing data under the null hypothesis.
Exploratory Data Analysis Techniques
Threshold for rejecting the null hypothesis, often 0.05.
Exploratory Data Analysis Techniques
Range likely to contain the true population parameter.
Exploratory Data Analysis Techniques
Compares means of two groups.
Exploratory Data Analysis Techniques
P-value: 'P' is for 'Probability' of the Null being true. If P is low, Null must go!
Exploratory Data Analysis Techniques
The exam often tests your ability to choose the correct statistical test for a given scenario. Pay attention to the number of groups being compared, the type of data (continuous vs. categorical), and whether samples are independent or paired. Keywords like 'compare means', 'association', or 'difference' are critical.
Exploratory Data Analysis Techniques
Confusing statistical significance with practical significance: A result can be statistically significant but too small to be practically important.
Exploratory Data Analysis Techniques
Incorrectly interpreting a high p-value: Failing to reject the null hypothesis does not mean the null hypothesis is true; it just means there isn't enough evidence to reject it.
Exploratory Data Analysis Techniques
Choosing the wrong statistical test: Using a t-test for categorical data or ANOVA for only two groups can lead to incorrect conclusions.
Exploratory Data Analysis Techniques
Symmetric, bell-shaped data distribution where mean, median, and mode are equal.
Exploratory Data Analysis Techniques
Measure of asymmetry in a data distribution, indicating a longer tail on one side.
Exploratory Data Analysis Techniques
Statistical measure quantifying the strength and direction of a linear relationship.
Exploratory Data Analysis Techniques
Measures linear relationship between two continuous variables (-1 to +1).
Exploratory Data Analysis Techniques
Measures monotonic relationship between ranked variables, useful for non-linear data.
Exploratory Data Analysis Techniques
Table showing correlation coefficients between multiple variables in a dataset.
Exploratory Data Analysis Techniques
Bar chart showing frequency distribution of numerical data in bins.
Exploratory Data Analysis Techniques
P-S-C: Pearson for Straight lines, Spearman for Curves (monotonic), Correlation is not Causation!
Exploratory Data Analysis Techniques
The exam often tests your ability to interpret correlation coefficients and understand the difference between correlation and causation. Be ready to distinguish between Pearson and Spearman based on data type and relationship linearity.
Exploratory Data Analysis Techniques
Assuming causation from correlation alone.
Exploratory Data Analysis Techniques
Using Pearson correlation for non-linear or ordinal relationships.
Exploratory Data Analysis Techniques
Ignoring outliers when analyzing distributions, as they can heavily skew results.
Exploratory Data Analysis Techniques
Learning from labeled data to predict outcomes.
Machine Learning Modeling Deep Dive
Finding patterns in unlabeled data.
Machine Learning Modeling Deep Dive
Agent learns by interacting with environment to maximize reward.
Machine Learning Modeling Deep Dive
Predicting a categorical label or class.
Machine Learning Modeling Deep Dive
Predicting a continuous numerical value.
Machine Learning Modeling Deep Dive
Grouping similar data points together.
Machine Learning Modeling Deep Dive
Measures difference between predictions and actual values.
Machine Learning Modeling Deep Dive
Algorithm to adjust model parameters to minimize loss.
Machine Learning Modeling Deep Dive
To remember the three ML paradigms: 'SUR'f the web to learn! S-Supervised, U-Unsupervised, R-Reinforcement.
Machine Learning Modeling Deep Dive
For the AWS ML Specialty exam, pay close attention to which AWS services align with each ML paradigm (e.g., Amazon Rekognition for supervised classification, Amazon SageMaker built-in algorithms like K-Means for unsupervised clustering). Understand the core difference between classification and regression.
Machine Learning Modeling Deep Dive
Using a regression algorithm for a classification problem (e.g., predicting 'spam' or 'not spam' with linear regression).
Machine Learning Modeling Deep Dive
Skipping data cleaning and preparation, leading to 'garbage in, garbage out' model performance.
Machine Learning Modeling Deep Dive
Applying a complex deep learning model when a simpler, more interpretable model would suffice and perform similarly.
Machine Learning Modeling Deep Dive
External configuration for an ML algorithm, set before training.
Machine Learning Modeling Deep Dive
Internal variable learned by the model from training data.
Machine Learning Modeling Deep Dive
Systematic exploration of all hyperparameter combinations.
Machine Learning Modeling Deep Dive
Samples hyperparameter values from defined distributions.
Machine Learning Modeling Deep Dive
Intelligent search strategy using a probabilistic model.
Machine Learning Modeling Deep Dive
Hyperparameter controlling step size in optimizer updates.
Machine Learning Modeling Deep Dive
Hyperparameter to prevent overfitting by penalizing complexity.
Machine Learning Modeling Deep Dive
AWS service for automated hyperparameter optimization.
Machine Learning Modeling Deep Dive
Hyperparameters are like the 'chef's recipe' (you set them before cooking), while model parameters are the 'flavors the food develops' (learned during cooking).
Machine Learning Modeling Deep Dive
The exam often tests your understanding of the difference between model parameters and hyperparameters. Remember that hyperparameters are set BEFORE training, while parameters are LEARNED DURING training. Also, be prepared to identify which AWS service helps automate hyperparameter tuning (SageMaker Automatic Model Tuning).
Machine Learning Modeling Deep Dive
Confusing model parameters with hyperparameters: Parameters are learned, hyperparameters are set.
Machine Learning Modeling Deep Dive
Using manual tuning for complex models: This is inefficient and often leads to suboptimal results.
Machine Learning Modeling Deep Dive
Ignoring the computational cost of tuning: Always consider strategies like early stopping or random search to save resources.
Machine Learning Modeling Deep Dive
Proportion of correct predictions among total predictions.
Machine Learning Modeling Deep Dive
Ratio of true positives to all positive predictions.
Machine Learning Modeling Deep Dive
Ratio of true positives to all actual positive instances.
Machine Learning Modeling Deep Dive
Harmonic mean of precision and recall, balancing both.
Machine Learning Modeling Deep Dive
Area under the Receiver Operating Characteristic curve, for classification.
Machine Learning Modeling Deep Dive
Mean Absolute Error; average absolute difference in regression.
Machine Learning Modeling Deep Dive
Root Mean Squared Error; square root of MSE, in original units.
Machine Learning Modeling Deep Dive
Splits data into K folds, trains K models, averages performance.
Machine Learning Modeling Deep Dive
PR-F1-ROC for Classification (People Really Feel 1st Round Of Coffee) and MAE-MSE-RMSE for Regression (My Aunt Eats Many Small Red Meals).
Machine Learning Modeling Deep Dive
For the exam, be able to distinguish between classification and regression metrics. Know that for imbalanced datasets, Accuracy is often a poor metric, and Precision, Recall, F1-score, and ROC AUC are preferred for classification. For regression, MAE, MSE, and RMSE are key. Understand that cross-validation is used to obtain a more robust estimate of model performance and prevent overfitting.
Machine Learning Modeling Deep Dive
Using Accuracy as the sole metric for highly imbalanced classification datasets.
Machine Learning Modeling Deep Dive
Not using a validation strategy like cross-validation, leading to overfitting assessment.
Machine Learning Modeling Deep Dive
Confusing classification metrics with regression metrics.
Machine Learning Modeling Deep Dive
Systematic errors in ML outputs leading to unfair outcomes.
Machine Learning Modeling Deep Dive
Quantitative measures to assess equitable model performance across groups.
Machine Learning Modeling Deep Dive
Equal positive prediction rates across different groups.
Machine Learning Modeling Deep Dive
Equal true positive rates across different groups.
Machine Learning Modeling Deep Dive
Equal true positive and false positive rates across groups.
Machine Learning Modeling Deep Dive
AWS service for detecting bias and explaining ML model predictions.
Machine Learning Modeling Deep Dive
Changes in model fairness over time, requiring continuous monitoring.
Machine Learning Modeling Deep Dive
To remember fairness metrics: 'D-E-E-P' for Demographic Parity, Equal Opportunity, Equalized Odds, and Predictive Parity. Each one measures a slightly different aspect of fairness.
Machine Learning Modeling Deep Dive
The exam frequently tests on the capabilities of AWS SageMaker Clarify for both bias detection and model explainability. Understand the difference between pre-training and post-training bias detection and the various fairness metrics it supports.
Machine Learning Modeling Deep Dive
Assuming overall high accuracy means the model is fair for all subgroups.
Machine Learning Modeling Deep Dive
Only checking for bias once during development and not monitoring in production.
Machine Learning Modeling Deep Dive
Believing there is a single 'fairness metric' that applies to all use cases.
Machine Learning Modeling Deep Dive
Making a trained ML model available for predictions.
ML Implementation and Operations (MLOps)
Generating predictions immediately for individual requests.
ML Implementation and Operations (MLOps)
Processing large datasets for predictions at once.
ML Implementation and Operations (MLOps)
Reducing precision of model weights for smaller size and faster inference.
ML Implementation and Operations (MLOps)
Removing less important connections in a neural network.
ML Implementation and Operations (MLOps)
A hosted, scalable real-time inference service for ML models.
ML Implementation and Operations (MLOps)
A service for processing large datasets with ML models.
ML Implementation and Operations (MLOps)
AWS-designed ML inference chip for high performance.
ML Implementation and Operations (MLOps)
To remember inference optimization techniques: 'C.H.A.P.': Caching, Hardware acceleration, Quantization, Pruning. These are your tools for speed!
ML Implementation and Operations (MLOps)
The exam frequently tests your knowledge of SageMaker's deployment options. Be prepared to distinguish between SageMaker Endpoints (real-time, low latency, high availability) and SageMaker Batch Transform (high throughput, offline processing). Keywords like 'immediate predictions' or 'large dataset processing' are key indicators.
ML Implementation and Operations (MLOps)
Deploying a large, unoptimized model for real-time inference, leading to high latency and costs.
ML Implementation and Operations (MLOps)
Using real-time endpoints for batch processing tasks, which is inefficient and expensive.
ML Implementation and Operations (MLOps)
Not considering model versioning and A/B testing during deployment, making updates risky.
ML Implementation and Operations (MLOps)
Statistical properties of input features change.
ML Implementation and Operations (MLOps)
Relationship between input features and target changes.
ML Implementation and Operations (MLOps)
Updating a model with new data to improve performance.
ML Implementation and Operations (MLOps)
AWS service for detecting data and model quality issues.
ML Implementation and Operations (MLOps)
Reference data/metrics for comparison in monitoring.
ML Implementation and Operations (MLOps)
Time taken for a model to make a prediction.
ML Implementation and Operations (MLOps)
DRIFT: Detect, Respond, Investigate, Fix, Train. A cycle for managing model degradation.
ML Implementation and Operations (MLOps)
The exam frequently tests your understanding of different types of drift (data vs. concept) and the appropriate AWS services for detecting and responding to them. Keywords like 'degradation,' 'staleness,' 'drift,' and 'retrain' are common.
ML Implementation and Operations (MLOps)
Ignoring operational metrics: Focusing only on model accuracy and overlooking latency or error rates can lead to production issues even with a 'good' model.
ML Implementation and Operations (MLOps)
One-size-fits-all retraining: Applying the same retraining schedule to all models, regardless of their sensitivity to drift or data update frequency, is inefficient.
ML Implementation and Operations (MLOps)
Not establishing a baseline: Without a clear baseline for data and model performance, it's impossible to accurately detect and quantify drift.
ML Implementation and Operations (MLOps)
Comparing two versions of a model in production to find the better performer.
ML Implementation and Operations (MLOps)
Continuous Integration/Continuous Delivery; automates software development and deployment.
ML Implementation and Operations (MLOps)
AWS service to orchestrate and automate ML workflows.
ML Implementation and Operations (MLOps)
A central repository for tracking and managing ML model versions and metadata.
ML Implementation and Operations (MLOps)
Gradually rolling out a new model version to a small subset of users.
ML Implementation and Operations (MLOps)
Think 'A/B Test for Better Models, CI/CD for Continuous Delivery.' The 'A' and 'B' remind you of comparing two options, while 'CI/CD' is about making everything flow smoothly and automatically.
ML Implementation and Operations (MLOps)
The exam frequently tests your understanding of how AWS services map to MLOps stages. Be ready to identify which service (e.g., SageMaker Pipelines, CodePipeline, SageMaker Model Registry) is used for specific tasks like workflow orchestration, model versioning, or automated deployment.
ML Implementation and Operations (MLOps)
Not defining clear success metrics before starting an A/B test, leading to ambiguous results.
ML Implementation and Operations (MLOps)
Ignoring data and model versioning in CI/CD pipelines, making reproducibility impossible.
ML Implementation and Operations (MLOps)
Failing to monitor models post-deployment, missing performance degradation or drift.
ML Implementation and Operations (MLOps)
Automated retraining of ML models on new data to maintain performance.
ML Implementation and Operations (MLOps)
Malicious input designed to trick or compromise an ML model.
ML Implementation and Operations (MLOps)
Process of removing personally identifiable information from datasets.
ML Implementation and Operations (MLOps)
Ability to recreate ML experiments and deployments consistently.
ML Implementation and Operations (MLOps)
MLOps is like a 'MODEL FACTORY': M-onitoring, O-perations, D-ata, E-ngineering, L-ifecycle. It's all about making sure your models are built, deployed, and maintained like a well-oiled machine!
ML Implementation and Operations (MLOps)
The exam often tests your understanding of the differences between MLOps and traditional DevOps, especially concerning data and model lifecycle management. Be prepared to identify AWS services that support each stage of a secure MLOps pipeline, such as S3 for data, SageMaker for training/deployment, and IAM/KMS for security.
ML Implementation and Operations (MLOps)
Underestimating the importance of data versioning and validation in MLOps.
ML Implementation and Operations (MLOps)
Neglecting continuous monitoring of model performance and data quality in production.
ML Implementation and Operations (MLOps)
Failing to implement robust access controls and encryption for sensitive ML data and models.
ML Implementation and Operations (MLOps)