AWS Certified Machine Learning – Specialty flashcards
144 free flashcards. Tap a card to flip it.
Secure On-Premises to S3 Ingestion
Flip cardUtilizing AWS services like DataSync, KMS, and CloudTrail to securely, scalably, and audibly transfer large volumes of sensitive data from on-premises environments to an S3 data lake.
- AWS DataSync: Optimized, secure data transfer from on-premises.
- AWS KMS: Manages encryption keys for data at rest in S3.
- AWS CloudTrail: Provides activity logs for auditing data access and transfers.
Memory trick: DataSync moves it, KMS encrypts it, CloudTrail records it!
Median Imputation for Skewed Data
Flip cardReplacing missing numerical values with the median of the existing values in that column, particularly effective for skewed distributions to maintain data integrity.
- Robust to outliers, unlike the mean.
- Preserves the shape of skewed distributions better.
- Suitable for continuous numerical features with missing values.
Memory trick: Missing data? Medians for skewed, Means for normal, Modes for categories!
Chi-squared Test of Independence
Flip cardA statistical test used to determine if there is a significant association between two categorical variables.
- Compares observed frequencies to expected frequencies.
- Null hypothesis: no association between variables.
- Alternative hypothesis: there is an association.
- Requires categorical data.
Memory trick: Categorical connections call for Chi-squared.
Seasonal Time-Series Imputation
Flip cardImputing missing values in time-series data by leveraging and preserving inherent seasonal or cyclical patterns (e.g., daily, weekly, monthly cycles).
- Accounts for recurring patterns in data.
- More sophisticated than simple statistical imputations.
- Methods include seasonal averages, seasonal Kalman filters, Prophet models.
- Crucial for accurate modeling of cyclical time series.
Memory trick: To fill gaps in cycles, decompose and impute for seasonal smiles.
SSE-KMS
Flip cardServer-Side Encryption with AWS Key Management Service (AWS KMS) managed keys encrypts objects in Amazon S3 using KMS keys. It offers an audit trail of key usage and allows customers to define key access policies.
- Encrypts data at rest in S3.
- Keys are managed within AWS KMS, providing customer control and auditing.
- AWS manages the encryption and decryption process.
Memory trick: KMS Keeps My Secrets Securely Stored.
Latent Dirichlet Allocation (LDA)
Flip cardAn unsupervised generative probabilistic model used for topic modeling, which presumes that documents are combinations of topics and that topics are combinations of words.
- Discovers abstract 'topics' from a collection of documents.
- Documents are modeled as mixtures of topics.
- Topics are modeled as mixtures of words.
- Requires specifying the number of topics (K) beforehand.
Memory trick: LDA lets documents discover their latent topics.
Logical Inconsistency Resolution
Flip cardA data cleaning strategy to address records where values across multiple features contradict logical or domain-specific rules, often by using predictive imputation or rule-based correction to ensure internal consistency.
- Involves multiple features with conflicting values.
- Aims to preserve data while correcting logical errors.
- Often requires domain expertise and predictive modeling for imputation.
Memory trick: Contradictory data's a tough call, predictive imputation saves it all.
Time-Series Imputation (Forward Fill)
Flip cardA technique for handling missing values in time-series data by carrying forward the last valid observation to fill subsequent missing data points.
- Preserves the temporal order of data.
- Suitable for data collected at irregular intervals.
- Simple to implement and computationally efficient.
- Can be extended with more sophisticated methods like interpolation.
Memory trick: For time series, just carry the last value forward, it's logical.
IQR Outlier Rule
Flip cardA statistical method for identifying outliers based on the interquartile range (IQR), where values falling below Q1 - 1.5 * IQR or above Q3 + 1.5 * IQR are considered outliers.
- Robust to skewed distributions and extreme values.
- Does not assume a specific data distribution (non-parametric).
- Commonly visualized using box plots.
Memory trick: IQR is the range, for outliers it's the gauge.
Image Preprocessing for CNNs
Flip cardStandardization of image dimensions and characteristics (e.g., resizing, cropping, normalization) to prepare them as consistent input for convolutional neural networks (CNNs).
- Consistent input size is crucial for CNNs.
- Maintaining aspect ratio prevents distortion.
- Common operations: resize, crop, pad.
- Normalization often follows to scale pixel values.
Memory trick: Shortest side first, then center crop, keeps the image looking sharp.
Power Transformation
Flip cardA family of transformations (e.g., Box-Cox, Yeo-Johnson) used to transform non-normally distributed data to a more Gaussian or symmetric distribution.
- Aims to stabilize variance and normalize data.
- Box-Cox requires strictly positive data.
- Yeo-Johnson can handle positive, negative, or zero values.
- Finds an optimal lambda (λ) parameter for the transformation.
Memory trick: Power transforms pinpoint the perfect push to normality.
Word Cloud
Flip cardA visual representation of text data where the size of each word indicates its frequency or importance within a given text corpus.
- Quickly highlights most frequent terms.
- Useful for text data EDA.
- Size and color can be used to convey additional information.
Memory trick: Words form clouds, showing what's most loud.
AWS Glue Streaming ETL
Flip cardA serverless data integration service that provides streaming ETL capabilities, built on Apache Spark Streaming, for transforming data from various sources to targets, handling schema evolution and complex data structures.
- Serverless Apache Spark Streaming environment.
- Handles schema inference and evolution.
- Ideal for flattening nested data and format conversions.
- Cost-effective and scalable for continuous data processing.
Memory trick: Glue streams and transforms, Spark-powered and serverless!
S3 Glacier Deep Archive
Flip cardThe lowest-cost Amazon S3 storage class for long-term archiving of data that is accessed once or twice a year, with retrieval times in hours.
- Lowest storage cost (per GB).
- Designed for archival with infrequent access.
- Retrieval times range from hours to 12 hours.
- High durability across multiple Availability Zones.
Memory trick: S3 Storage: Standard for hot, IA for warm, Glacier for cold!
AWS Glue for Data Transformation
Flip cardAWS Glue is a serverless data integration service that facilitates ETL (Extract, Transform, Load) operations, including data cleaning, transformation, and preparation for analytics and machine learning.
- Serverless and scalable
- Apache Spark-based for distributed processing
- Includes Glue Data Catalog for metadata management
- Supports various data sources and targets
Memory trick: To clean up text data at scale, remember Glue is your best ally.
Chi-squared Goodness-of-Fit Test
Flip cardA statistical test used to determine if an observed frequency distribution for a single categorical variable differs significantly from an expected theoretical distribution.
- Compares observed frequencies to expected frequencies.
- Null hypothesis: observed distribution fits the expected distribution.
- Alternative hypothesis: observed distribution does not fit the expected distribution.
- Requires categorical data and expected frequencies for each category.
Memory trick: Chi-squared goodness-of-fit checks if the observed fits the 'should'.
AWS Glue Data Catalog
Flip cardA fully managed, persistent metadata store that acts as a central repository for schema, location, and partition information of data assets in a data lake.
- Serverless metadata management
- Integrates with Athena, Glue ETL, Redshift Spectrum
- Supports various data formats (CSV, JSON, Parquet)
- Enables unified view of data lake data
Memory trick: Glue's catalog is the map to your data lake's schema trap.
PII Masking with AWS Glue
Flip cardUsing AWS Glue ETL jobs to transform sensitive PII data in datasets by replacing it with anonymized, pseudonymized, or fictitious values, often maintaining referential integrity.
- AWS Glue provides a serverless Spark-based environment.
- Custom Python/PySpark scripts enable flexible masking logic.
- Suitable for large-scale structured data in S3.
- Can implement consistent hashing or tokenization for referential integrity.
Memory trick: Glue's custom scripts make PII disappear, leaving consistent stand-ins clear.
Time-Series Anomaly Detection (Moving Average)
Flip cardA technique for identifying anomalous data points in time-series data by comparing each point to a statistically derived 'normal' range (e.g., mean ± std dev) calculated from a preceding moving window of data.
- Considers temporal context (local patterns).
- Effective for detecting sudden spikes or drops.
- Adapts to changing baseline trends over time.
Memory trick: Moving average sees the local flow, spikes above it must go.
2D Kernel Density Estimate (KDE)
Flip cardA non-parametric way to estimate the probability density function of two continuous random variables, often visualized as contours or a heatmap on a scatter plot, showing areas of higher data concentration.
- Visualizes joint distribution of two continuous variables.
- Does not assume a specific distribution shape.
- Useful for identifying clusters and density variations.
Memory trick: Scatter points tell the story, KDE shows the density's glory.
Real-time IoT Data Pipeline
Flip cardA system for ingesting, processing, and storing high-volume, low-latency data from IoT devices using serverless AWS services.
- AWS IoT Core for device connectivity and message broker.
- Amazon Kinesis Data Firehose for serverless ingestion, buffering, and delivery.
- Amazon S3 as the scalable and cost-effective data lake storage.
- Supports basic transformations and data format conversions.
Memory trick: IoT Core connects, Firehose streams, S3 stores the data dreams.
Handling Impossible Values
Flip cardImpossible values are data points that contradict known constraints or physical laws, such as negative age or a percentage greater than 100%.
- Indicate data corruption or entry errors.
- Should generally be removed or corrected to ensure data integrity.
- Different from outliers, which are extreme but plausible values.
Memory trick: Impossible values - just toss them out!
S3 Data Security Best Practices
Flip cardA combination of encryption, access control, and secure communication protocols to protect data stored in Amazon S3 for machine learning workloads.
- Encryption at rest: Use Server-Side Encryption with KMS (SSE-KMS) for managed keys and auditability.
- Encryption in transit: Enforce HTTPS for all data transfer to and from S3.
- Access control: Implement granular permissions using IAM policies and S3 bucket policies.
Memory trick: Secure S3, Keep Data Safe, IAM and KMS are the Keys!
Real-time Serverless Data Ingestion & Processing
Flip cardA common AWS architectural pattern for handling high-volume, continuous data streams, transforming them in near real-time, and storing them for machine learning workloads.
- Uses Kinesis Data Streams for data ingestion.
- Leverages AWS Lambda for serverless, event-driven processing.
- Stores processed data in Amazon S3 for data lake capabilities.
- Scalable, cost-effective, and fully managed components.
Memory trick: Stream with Kinesis, Lambda processes, S3 stores the ML insights.
Hierarchical Clustering
Flip cardAn unsupervised clustering algorithm that builds a hierarchy of clusters, either by starting with individual data points and merging them (agglomerative) or by starting with one large cluster and splitting it (divisive).
- Does not require pre-specifying the number of clusters.
- Results are often visualized with a dendrogram.
- Useful for exploring natural groupings in data.
- Can be computationally intensive for very large datasets.
Memory trick: Hierarchical clustering helps you harvest clusters from a growing tree.
S3 Data Optimization for Analytics
Flip cardStrategies to enhance query performance and reduce costs for analytics on data stored in Amazon S3, typically involving columnar file formats and effective partitioning.
- Columnar formats (Parquet, ORC) reduce I/O by reading only required columns.
- Partitioning reduces data scanned by allowing query engines to skip irrelevant S3 prefixes.
- Compression reduces storage costs and improves query speed.
- Choosing partition keys based on common query filters is crucial.
Memory trick: Parquet for columns, dates for partitions, makes queries sing with no hesitations.
Pseudonymization (Data Masking)
Flip cardA data privacy technique where direct identifiers are replaced with artificial identifiers, ensuring individuals cannot be identified without additional information, while preserving data utility.
- Key for compliance (e.g., HIPAA, GDPR).
- Replaces sensitive data with non-sensitive substitutes.
- Can maintain format and referential integrity.
- Differs from anonymization (which aims for irreversible de-identification).
Memory trick: Masking hides the identities, keeping data safe and useful.
Log Transformation with Constant
Flip cardA data transformation technique (log(x+c)) used to reduce positive skewness in a numerical feature, where a small positive constant (c) is added to each value to allow transformation of zero or negative values.
- Effective for right-skewed data.
- Handles zero values, converting them to log(c).
- Preserves order and relative relationships.
Memory trick: Log(x+c) makes skewed data agree, especially when zero's in the tree.
S3 Storage Class Optimization
Flip cardSelecting the most appropriate Amazon S3 storage class for data based on access patterns, durability requirements, and cost objectives to optimize storage spend.
- S3 Standard: Frequent access, high performance.
- S3 Standard-IA: Infrequent access, rapid retrieval.
- S3 Intelligent-Tiering: Automatic tiering for changing/unknown access.
- S3 Glacier Flexible Retrieval: Archival, flexible retrieval options.
Memory trick: Match your data's 'hotness' to S3's 'coolness' for optimal savings.
AWS IoT Core for Device Ingestion
Flip cardA managed cloud platform that lets connected devices (IoT) easily and securely interact with cloud applications and other devices, handling device authentication, message ingestion, and routing.
- Connects and manages billions of IoT devices.
- Supports various device protocols (MQTT, HTTPS, WebSockets).
- Ingests small, high-frequency messages efficiently.
- Routes messages to other AWS services (S3, Kinesis, Lambda).
Memory trick: IoT Core connects your devices, streams their data to the cloud.
Median Imputation
Flip cardA data imputation technique where missing values in a feature are replaced by the median value of the non-missing entries in that feature.
- Robust to outliers and skewed distributions.
- Preserves the original distribution shape better than mean imputation for skewed data.
- Suitable for numerical data.
Memory trick: Don't 'mean' to skew, 'median' is the clean way through.
Amazon S3 for ML Data Lakes
Flip cardAmazon S3 (Simple Storage Service) is an object storage service offering industry-leading scalability, data availability, security, and performance. It's often used as the foundation for data lakes for machine learning.
- Stores any type of object (data lake foundation).
- High durability (11 nines) and availability.
- Encryption at rest (SSE-S3, SSE-KMS, SSE-C) and in transit (SSL/TLS).
- Highly scalable, cost-effective for large datasets.
Memory trick: Securely store your ML treasures in the scalable S3 cloud.
Independent Samples T-test
Flip cardA statistical hypothesis test used to determine if there is a significant difference between the means of two independent groups.
- Compares means of two unrelated samples.
- Assumes normality and equal variances (or uses Welch's t-test if variances are unequal).
- Null hypothesis: means are equal.
Memory trick: T-test for two, ANOVA for many groups to view.
Amazon Kinesis Data Firehose
Flip cardA fully managed service for delivering real-time streaming data to destinations like Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and Splunk, with built-in buffering and transformation.
- Fully managed and automatically scales.
- No servers to provision or manage.
- Buffering and batching capabilities.
- Supports basic transformations (e.g., format conversion, compression).
Memory trick: Firehose funnels data streams to storage, fast and easy.
Frequency Encoding
Flip cardA categorical feature encoding technique where each category is replaced by the frequency or count of its occurrence in the dataset. It's useful for high-cardinality features to reduce dimensionality.
- Reduces dimensionality for high-cardinality categorical features
- Preserves information about category prevalence
- Does not introduce arbitrary ordinality
- Can be effective if frequency correlates with the target
Memory trick: When too many categories make a mess, frequency encoding is your best address.
Multiple Imputation by Chained Equations (MICE)
Flip cardA sophisticated imputation technique that iteratively imputes missing data using a series of regression models, creating multiple complete datasets to account for imputation uncertainty.
- Handles Missing At Random (MAR) and Missing Completely At Random (MCAR) data.
- Preserves variability and relationships better than single imputation.
- Generates multiple imputed datasets, then pools results.
- Computationally more intensive but provides more accurate inferences.
Memory trick: MICE makes multiple models to make missing data meaningful.
AWS Data Governance with Lake Formation
Flip cardEstablishing and enforcing policies for data access, usage, and auditing within a data lake, primarily using AWS Lake Formation for centralized control.
- AWS Lake Formation for centralized, granular access control.
- Integrates with AWS Glue Data Catalog for metadata.
- Supports table-, column-, and row-level permissions.
- Provides approval workflows for data access.
Memory trick: Lake Formation governs, Glue catalogs, CloudTrail audits all data logs.
Pearson Correlation Coefficient
Flip cardA measure of the linear correlation between two continuous variables, ranging from -1 (perfect negative linear correlation) to +1 (perfect positive linear correlation), with 0 indicating no linear correlation.
- Measures linear relationships only.
- Values range from -1 to +1.
- Sensitive to outliers.
Memory trick: Pearson's 'R' for linear 'relations' you see.
Amazon SageMaker Processing Jobs
Flip cardA fully managed service within Amazon SageMaker for running data processing, feature engineering, data validation, and model evaluation workloads.
- Managed, scalable environment for custom scripts.
- Supports various frameworks (Scikit-learn, Spark, custom containers).
- Ideal for large-scale data preparation before model training.
- Automates infrastructure provisioning and resource management.
Memory trick: SageMaker's Process Jobs transforms images for better models.
Hybrid Recommendation Systems
Flip cardHybrid recommendation systems combine multiple recommendation approaches (e.g., collaborative filtering, content-based filtering) to leverage their respective strengths and mitigate weaknesses, such as the cold start problem.
- Address cold start by using content-based features for new users/items.
- Can improve overall recommendation quality and diversity.
- Common combination: matrix factorization (collaborative) + content-based.
- Different integration strategies: weighted, switching, mixed, or cascade.
Memory trick: For 'NEWBIES', don't be cold, 'MIX' your methods to warm them up!
Recall vs. Precision Trade-off
Flip cardRecall measures the proportion of actual positives correctly identified (minimizing false negatives), while Precision measures the proportion of predicted positives that are actually correct (minimizing false positives). There is often a trade-off between these two metrics.
- High Recall is crucial when false negatives are expensive (e.g., fraud, disease detection).
- High Precision is crucial when false positives are expensive (e.g., spam detection, recommending wrong products).
- F1-score is a harmonic mean used when both are equally important.
- The business context dictates which metric to prioritize.
Memory trick: When 'COSTS ARE HIGH', choose your 'METRIC WISELY' to match the most critical error.
Dropout
Flip cardDropout is a regularization technique used in neural networks where a random subset of neurons is temporarily ignored (dropped out) during each training iteration.
- Forces the network to learn more robust features.
- Prevents complex co-adaptations between neurons.
- Acts as an ensemble of many smaller networks, reducing overfitting.
Memory trick: When neurons 'drop out', they can't 'gossip' and 'overfit' the training data.
Popularity Bias
Flip cardA bias in recommendation systems where items that are already popular tend to be recommended more frequently, leading to a lack of diversity and potentially ignoring niche preferences.
- Results in 'rich-get-richer' phenomenon for popular items.
- Reduces discovery of long-tail items.
- Can lead to user dissatisfaction and filter bubbles.
Memory trick: Popularity's Pull, Niche's Fall.
Recall (Sensitivity)
Flip cardThe proportion of actual positive instances that were correctly identified by the model. It measures the model's ability to find all positive samples.
- Calculated as True Positives / (True Positives + False Negatives).
- Crucial for imbalanced datasets where missing positive cases is costly.
- Also known as sensitivity or True Positive Rate.
Memory trick: When positives are rare, recall's the flair, to catch every failure, you must be aware.
Model Simplicity for Small Data
Flip cardThe preference for simpler machine learning models (e.g., SVM, Logistic Regression) over complex ones (e.g., Deep Neural Networks) when dealing with small datasets to prevent overfitting and ensure better generalization.
- Complex models require more data to learn robust patterns.
- Simple models have lower variance, making them less prone to overfitting.
- SVMs are particularly effective due to margin maximization.
Memory trick: Small data, big problem, simple model's the anthem, SVM's the champion, where complexity is banned.
Environmental Robustness Testing
Flip cardThe process of evaluating a machine learning model's performance and stability when exposed to various natural, real-world variations or perturbations in its input data (e.g., changes in lighting, rotation, noise, blur).
- Ensures model generalizes well to diverse real-world conditions.
- Tests resilience to expected, non-malicious input variations.
- Crucial for safety-critical applications.
Memory trick: Environmental robustness, like a sturdy tree, withstands the wind, the rain, for all to see.
XGBoost for Imbalanced Data
Flip cardXGBoost is an optimized distributed gradient boosting library designed to be highly efficient, flexible, and portable. It excels in handling imbalanced datasets due to built-in mechanisms like `scale_pos_weight` and its robust ensemble nature.
- Gradient Boosting algorithm, builds trees sequentially.
- Parameters like `scale_pos_weight` directly address class imbalance.
- Provides feature importance scores for interpretability.
- Known for high performance and accuracy across various tasks.
Memory trick: When data isn't balanced, 'BOOST' your model's chances by weighting the rare cases.
Label Smoothing
Flip cardA regularization technique that replaces hard labels (0 or 1) with smoothed probability distributions during training, preventing the model from becoming overconfident and improving calibration.
- Reduces overconfidence in predictions.
- Improves model generalization.
- Adds a small amount of noise to the target labels.
Memory trick: Label smoothing, like a gentle hand, guides the model to understand, that certainty's a spectrum, not just a command.
Fine-tuning Pre-trained LMs
Flip cardThe process of adapting a large language model, pre-trained on a massive general corpus, to a specific downstream task or domain by continuing its training on a smaller, task-specific dataset.
- Leverages vast general knowledge from pre-training.
- Efficiently adapts to domain-specific language and tasks.
- Requires significantly less data than training from scratch.
Memory trick: Pre-trained models are like wise old sages, fine-tuning teaches them new tricks for specific stages.
Popularity Bias & Long-Tail Coverage
Flip cardPopularity bias in recommendation systems refers to the tendency of models to recommend disproportionately popular items, often overlooking niche or 'long-tail' items. Long-tail coverage is a metric that quantifies the percentage of items that are in the 'long tail' (less popular) that are actually recommended by the system.
- Common in collaborative filtering due to abundant data for popular items.
- Leads to lack of diversity and serendipity in recommendations.
- Metrics like Gini coefficient, catalog coverage, and long-tail coverage help detect it.
- Can be mitigated by re-ranking, boosting niche items, or hybrid models.
Memory trick: To see if your recommender is 'FAIR', check its 'DISTRIBUTION' and 'REACH'.
LIME (Local Interpretable Model-agnostic Explanations)
Flip cardA technique that explains the predictions of any machine learning model by approximating it locally with an interpretable model.
- Model-agnostic: Works with any black-box model.
- Provides local explanations: Explains a single prediction.
- Generates an interpretable model (e.g., linear model) around the instance of interest.
Memory trick: LIME explains local predictions, like zooming in.
Model Quantization
Flip cardModel quantization is a technique to reduce the precision of numbers used to represent a model's parameters and computations, typically from floating-point to lower-bit integers.
- Reduces model size significantly.
- Speeds up inference time due to simpler arithmetic operations.
- Can be applied during training (quantization-aware training) or post-training.
- May lead to a slight drop in accuracy, which needs to be balanced against performance gains.
Memory trick: Optimize models to make them 'LIGHT, FAST, and SMART' for deployment.
Stacking (Stacked Generalization)
Flip cardAn advanced ensemble technique where multiple base models are trained, and their predictions are then used as input features to a 'meta-learner' model, which learns to optimally combine these predictions to make the final output.
- Combines diverse models effectively.
- Meta-learner learns complex combination rules.
- Often yields higher performance than simple ensembles.
- Involves training multiple layers of models.
Memory trick: Stack the Models, Learn the Best Combine.
Log Transformation for Skewed Data
Flip cardLog transformation is a data preprocessing technique that applies the logarithm function to a numerical feature, commonly used to reduce skewness and stabilize variance in positively skewed distributions.
- Effective for features with a long tail towards larger values (positive skew).
- Helps make the data distribution more symmetrical, closer to normal.
- Beneficial for algorithms sensitive to feature distribution and scale.
Memory trick: When data is 'skewed', 'log' it to 'straighten' it out!
L1 Regularization (Lasso)
Flip cardL1 regularization, or Lasso regularization, adds a penalty term to the loss function that is proportional to the absolute value of the magnitude of the coefficients.
- Encourages sparsity in weights, effectively performing feature selection.
- Can drive some feature weights to exactly zero, simplifying the model.
- Helps mitigate overfitting by reducing model complexity.
Memory trick: To stop overfitting, 'Lasso' the complexity before it 'runs away' with your data.
SMOTE
Flip cardSynthetic Minority Over-sampling Technique, an oversampling method that generates synthetic samples for the minority class by interpolating between existing minority class instances and their nearest neighbors.
- Addresses class imbalance by increasing minority class size.
- Creates synthetic, not duplicate, samples.
- Helps prevent overfitting compared to simple oversampling.
Memory trick: SMOTE's the antidote, for tiny classes, it's the vote, creating new friends, so the model can quote.
Knowledge Distillation
Flip cardA model compression technique where a smaller 'student' model learns to reproduce the output probabilities (soft targets) of a larger, more complex 'teacher' model.
- Enables deployment of smaller, faster models.
- Student model often performs better than if trained directly on hard targets.
- Reduces computational cost and memory footprint.
Memory trick: For speed, distill knowledge into a smaller model.
Bias-Variance Trade-off
Flip cardThe dilemma in supervised learning where reducing one type of error (bias or variance) often increases the other.
- High bias (underfitting) means the model is too simple.
- High variance (overfitting) means the model is too complex.
- The goal is to find a balance for optimal generalization.
Memory trick: Simple models are biased, complex models are varied.
Robustness Testing (Adversarial Examples)
Flip cardRobustness testing evaluates how well a machine learning model performs when its input data is subjected to small, often imperceptible, perturbations or noise. Adversarial examples are inputs specifically crafted to cause a model to make an incorrect prediction with high confidence.
- Measures model's resilience to input variations and attacks.
- Crucial for safety-critical applications (e.g., autonomous driving, medical).
- Involves generating adversarial examples or injecting noise.
- Helps identify vulnerabilities and improve model reliability.
Memory trick: To check if your model is 'TOUGH', hit it with 'SMALL CHANGES' and see if it breaks.
Recurrent Neural Network (RNN)
Flip cardA type of neural network designed to process sequential data by maintaining an internal state (memory) that captures information from previous elements in the sequence.
- Suitable for time series, natural language, and speech.
- Captures temporal dependencies.
- Suffers from vanishing/exploding gradients in vanilla form, often mitigated by LSTMs/GRUs.
Memory trick: RNNs remember the past, like a flowing river, each drop influencing the next.