AWS Certified AI Practitioner flashcards
100 free flashcards. Tap a card to flip it.
Generative Models
Flip cardA type of foundation model that learns the underlying patterns and distributions of data to generate new, realistic samples that resemble the training data.
- Creates new data instances
- Learns complex data distributions
- Used for synthetic data, content creation, style transfer
Memory trick: Generative models GENERATE new ideas, like a creative artist.
AWS CloudTrail
Flip cardAn AWS service that enables governance, compliance, operational auditing, and risk auditing of your AWS account by logging actions taken by a user, role, or an AWS service.
- Records API calls and events.
- Logs actions across AWS services.
- Essential for security analysis and compliance.
- Logs can be stored in S3 and analyzed with other tools.
Memory trick: CloudTrail leaves a trail of every action in the cloud.
Word Embeddings
Flip cardWord embeddings are dense vector representations of words that encode their semantic meaning and relationships. Words with similar meanings are mapped to nearby points in the vector space.
- Capture semantic and syntactic relationships.
- Lower-dimensional and denser than one-hot encodings.
- Learned from large text corpora (e.g., Word2Vec, GloVe, FastText).
Memory trick: Words to numbers, retaining meaning.
Removing Identifier Features
Flip cardA feature engineering technique where unique identification columns (e.g., IDs, serial numbers) are removed from the dataset before model training, as they typically do not contribute to learning generalizable patterns and can cause overfitting.
- Prevents models from memorizing specific instances.
- Reduces dimensionality and potential for overfitting.
- Applies to features like CustomerID, OrderID, SSN.
Memory trick: IDs are unique, not general, so remove them.
Label (Supervised Learning)
Flip cardIn supervised machine learning, a label is the target variable or the 'correct answer' that the model is trained to predict for a given input. For classification tasks, labels are discrete categories.
- The 'correct answer' for input.
- Target variable in supervised learning.
- Discrete categories for classification.
Memory trick: Labeled data teaches smart models.
Transformer Architecture
Flip cardA neural network architecture, primarily relying on self-attention mechanisms, designed to process sequential data and capture long-range dependencies efficiently.
- Uses self-attention to weigh the importance of different parts of the input sequence.
- Processes entire sequences in parallel, unlike RNNs.
- Forms the foundation for many state-of-the-art Generative AI models, especially LLMs.
Memory trick: Networks Learn, Architectures Define How.
Temperature Parameter (Generative AI)
Flip cardA hyperparameter used during the inference phase of generative AI models (especially language models) to control the randomness or 'creativity' of the generated output by scaling the logits before the softmax function.
- Higher temperature (e.g., >1.0) increases randomness, leading to more diverse/surprising outputs.
- Lower temperature (e.g., <1.0) makes outputs more deterministic and predictable, often closer to the training data distribution.
- A temperature of 0 makes the model greedy, always picking the most probable token.
Memory trick: Turn up the temperature for more creative tunes.
One-Hot Encoding
Flip cardA technique to convert categorical variables into a numerical format that machine learning algorithms can understand, by creating binary columns for each category.
- Used for nominal (unordered) categorical data.
- Creates N new binary features for N categories.
- Avoids implying an artificial ordinal relationship.
- Can lead to increased dimensionality ('curse of dimensionality').
Memory trick: Categories need numbers, but order matters.
Fidelity (Generative AI)
Flip cardA key evaluation criterion for generative AI models, referring to the quality, realism, and accuracy of the generated outputs, and their adherence to the underlying data distribution and real-world constraints.
- Measures how 'real' or high-quality the generated content is.
- Crucial for applications where outputs must be physically or semantically correct.
- Often balanced against diversity and novelty in model design.
Memory trick: Generative AI needs Fidelity, Diversity, and Novelty.
Recommendation Systems (ML Task)
Flip cardA machine learning task focused on predicting a user's preference for an item or suggesting items that a user might be interested in, based on historical data and patterns.
- Aims to personalize user experience and increase engagement/sales.
- Often uses collaborative filtering, content-based filtering, or hybrid approaches.
- Commonly seen in e-commerce, streaming services, and social media.
Memory trick: Classify, Regress, Cluster, Recommend to learn.
Unsupervised Learning (Anomaly Detection)
Flip cardA machine learning approach used to identify rare items, events, or observations that deviate significantly from the majority of the data, without requiring labeled examples of anomalies.
- Learns patterns from unlabeled data.
- Identifies outliers or deviations from learned normal behavior.
- Common algorithms include K-Means, Isolation Forest, One-Class SVM.
- Useful when anomalies are rare or undefined.
Memory trick: Learning types: teacher, no teacher, or rewards.
Generative Adversarial Networks (GANs)
Flip cardA class of generative AI models consisting of two neural networks, a generator and a discriminator, that compete against each other to produce new, realistic data.
- Generator creates synthetic data.
- Discriminator distinguishes real from synthetic data.
- Used for image generation, style transfer, and data augmentation.
Memory trick: Generative models 'create' new things, often through 'adversarial' competition.
Attention Mechanism (AI)
Flip cardA component in neural networks that allows the model to dynamically weigh the importance of different parts of the input sequence when processing or generating specific parts of the output sequence.
- Enables models to focus on relevant information.
- Crucial for handling long-range dependencies in sequences.
- Forms the core of Transformer architectures.
Memory trick: Attention pays attention to what matters.
Feature Engineering: Removing IDs
Flip cardUnique identifiers like customer IDs should generally be removed from datasets before training an AI/ML model because they do not contribute predictive information and can lead to issues like overfitting.
- IDs are unique, not predictive.
- Including IDs can cause overfitting.
- Remove IDs during data preprocessing.
Memory trick: Clean data for smarter decisions.
Classification (ML)
Flip cardA supervised machine learning task that involves predicting a discrete class label for a given input data point.
- Output is a category or class (e.g., spam/not spam, positive/negative sentiment).
- Requires labeled training data.
- Examples: sentiment analysis, image recognition, medical diagnosis.
Memory trick: ML 'tasks' are about 'predicting' different 'types' of outcomes.
Fine-tuning (Transfer Learning for LLMs)
Flip cardFine-tuning is a technique where a pre-trained large language model (LLM) is further trained on a smaller, domain-specific or task-specific dataset to adapt its knowledge and capabilities to a new context.
- Leverages pre-trained knowledge.
- Uses smaller, specialized datasets.
- More efficient than training from scratch.
Memory trick: Refine the giant for a niche.
Fine-tuning (LLMs)
Flip cardThe process of taking a pre-trained large language model and further training it on a smaller, domain-specific dataset to adapt its capabilities to a particular task or domain.
- Requires a base pre-trained model.
- Uses a smaller, specialized dataset.
- Adapts model to specific tasks/domains.
- More efficient than training from scratch.
Memory trick: LLMs 'learn' general, then get 'fine-tuned' for 'specific' expertise.
Hallucination (Generative AI)
Flip cardThe phenomenon where a generative AI model produces outputs that are plausible-sounding or grammatically correct but are factually incorrect, nonsensical, or not grounded in the input or training data.
- Often occurs when models lack sufficient information or are prompted with ambiguous queries.
- A significant challenge in ensuring the factual accuracy of generative AI outputs.
- Can be mitigated by techniques like retrieval augmented generation (RAG).
Memory trick: Generative models can hallucinate facts.
Feature Scaling
Flip cardFeature scaling is a data preprocessing technique used to standardize or normalize the range of independent numerical features in a dataset. This prevents features with larger values from dominating the model's learning process and helps optimization algorithms converge faster.
- Standardizes/normalizes numerical features.
- Prevents dominance by large values.
- Helps model convergence.
Memory trick: Clean and shape data for superior models.
Generative AI Core Function
Flip cardGenerative AI models learn the patterns and structure of input data to create new, original data that resembles the training data.
- Learns data distribution.
- Creates novel samples (text, images, audio, data).
- Distinguished from discriminative models which predict labels.
Memory trick: Generative AI 'creates' 'new' things from 'learned' patterns.
Unsupervised Learning for Anomaly Detection
Flip cardA machine learning approach where models learn patterns from unlabeled data, typically identifying anomalies as data points that deviate significantly from the learned 'normal' patterns.
- Does not require labeled anomaly data for training.
- Learns the underlying structure of the majority (normal) class.
- Effective for rare events or when labeling is difficult/expensive.
- Examples: Clustering, Density Estimation, Autoencoders.
Memory trick: ML 'learns' from data, either 'with' or 'without' a teacher.
Anomaly Detection (ML Task)
Flip cardAnomaly detection is a machine learning task focused on identifying rare observations or events that deviate significantly from the majority of the data, often indicating a problem or unusual activity.
- Identifies outliers/deviations.
- Used for fraud, error, or intrusion detection.
- Often involves unsupervised or semi-supervised learning.
Memory trick: Models predict, categorize, group, or find oddities.
Classification (ML Task)
Flip cardClassification is a supervised machine learning task where the model learns from labeled data to predict a discrete category or class for new, unseen data points.
- Predicts discrete, distinct categories (e.g., 'spam'/'not spam', 'cat'/'dog').
- Requires labeled training data.
- Can be binary (two classes) or multi-class (more than two classes).
Memory trick: Labeled data leads to predictable outcomes.
Transformer Architecture (Generative AI)
Flip cardThe Transformer architecture, particularly its decoder-only variants, is a neural network design that utilizes self-attention mechanisms to efficiently process and generate sequential data, excelling in tasks like text generation due to its ability to capture long-range dependencies and produce highly coherent outputs.
- Uses self-attention.
- Excels in sequential data (text).
- Captures long-range dependencies.
Memory trick: AI models build new realities.
Recurrent Neural Network (RNN)
Flip cardA class of artificial neural networks where connections between nodes form a directed graph along a temporal sequence, allowing them to exhibit temporal dynamic behavior.
- Designed for sequential data (e.g., text, speech, time series).
- Has an internal memory (hidden state) to retain information from previous inputs.
- Suffers from vanishing/exploding gradient problems in long sequences (addressed by LSTMs/GRUs).
Memory trick: Networks 'connect' neurons to 'learn' different 'patterns'.
Attention Mechanism (Generative AI)
Flip cardThe attention mechanism in generative AI models (especially Transformers) allows the model to selectively focus on and weigh the importance of different parts of the input sequence when generating an output, enabling effective handling of long-range dependencies and improved contextual understanding.
- Selectively focuses on input parts.
- Weighs importance of input elements.
- Crucial for long-range dependencies.
Memory trick: AI's memory for the long haul.
Precision
Flip cardA metric that measures the proportion of positive identifications that were actually correct. It is calculated as True Positives / (True Positives + False Positives).
- Focuses on minimizing False Positives (incorrectly identified positives).
- Important when the cost of a false alarm is high (e.g., flagging legitimate transactions as fraud, spam detection).
- Answers the question: 'Of all items predicted positive, how many are actually positive?'
Memory trick: PR-F1, Accuracy's good, but Precision avoids false alarms.
Inference (AI/ML)
Flip cardInference, in AI/ML, refers to the process of using a trained machine learning model to make predictions or generate outputs on new, unseen data, applying the knowledge it acquired during training.
- Also known as prediction or generation (for generative models).
- Occurs after the model has been trained and validated.
- The stage where the model is put into practical use.
Memory trick: Data-Train-Evaluate-Infer: The cycle of an AI.
ML Lifecycle: Training
Flip cardThe phase where a machine learning model learns patterns and relationships from the prepared data to make predictions or decisions.
- Occurs after data preparation.
- Involves feeding the algorithm with labeled or unlabeled data.
- The model adjusts its internal parameters based on the data.
Memory trick: Life 'cycles' through data, training, and deploying models.
Inference (AI/ML Lifecycle)
Flip cardThe process of using a trained machine learning model to make predictions or classifications on new, unseen data.
- Occurs after a model has been trained and validated.
- Applies the learned patterns to real-world data.
- The primary goal of deploying an ML model.
- Can be performed in batch or real-time.
Memory trick: ML cycle: data, build, test, use.
Transfer Learning (Fine-tuning LLMs)
Flip cardTransfer learning, specifically fine-tuning, is a technique where a pre-trained large language model (LLM) is further trained on a smaller, domain-specific dataset to adapt its knowledge and performance to a particular task or domain.
- Leverages general knowledge from pre-training.
- Requires less data and computational resources than training from scratch.
- Effective for specialized tasks or domains.
- Involves adjusting specific layers or the entire model.
Memory trick: General knowledge plus specific training equals expertise.
False Positive Rate (FPR)
Flip cardThe False Positive Rate (FPR), also known as the fallout, is a metric that quantifies the proportion of negative instances that were incorrectly classified as positive by a model.
- Calculated as False Positives / (False Positives + True Negatives).
- Indicates the rate of Type I errors.
- Relevant in scenarios where misclassifying a negative as positive is costly (e.g., false alarms).
Memory trick: Confusion Matrix: The four corners tell the story.
Deep Semantic Understanding & Reasoning (FMs)
Flip cardThe ability of foundation models to not just process surface-level information but to grasp the underlying meaning, context, and relationships within data, enabling complex inference and problem-solving.
- Goes beyond keyword matching to conceptual understanding.
- Crucial for tasks requiring logical deduction, problem-solving, and abstract thought.
- Often emerges with increased model size and training data diversity.
Memory trick: Deep reasoning unlocks scientific breakthroughs.
Foundation Model Use Case: Content Generation
Flip cardUtilizing foundation models, especially LLMs, to automatically create diverse forms of text content, such as marketing copy, articles, emails, or code.
- Relies on the model's ability to understand context and generate coherent, creative text.
- Can be personalized based on input parameters.
- Requires careful prompting and potentially fine-tuning for specific styles or tones.
Memory trick: FMs tackle many tasks, from creative content to deep analysis.
Temperature (Generative AI)
Flip cardA hyperparameter in generative AI models that controls the randomness or 'creativity' of the generated output by scaling the logit probabilities before sampling.
- Higher temperature (e.g., >1.0) increases randomness, leading to more diverse but potentially less coherent output.
- Lower temperature (e.g., <1.0) decreases randomness, leading to more predictable and focused output.
- A temperature of 0 makes the model fully deterministic, always picking the highest probability option.
- Applied during the inference phase of generation.
Memory trick: Temperature Tunes Output Creativity
On-premises Foundation Model Deployment
Flip cardRunning a foundation model entirely within an organization's own private data centers or private cloud infrastructure, without reliance on external cloud providers for model inference or data processing.
- Provides maximum control over data security, privacy, and sovereignty.
- Requires significant internal infrastructure and expertise.
- Can be costly due to hardware and operational overhead.
Memory trick: On-premises: Your castle, your rules, your data.
Generative Repetitiveness
Flip cardA common problem in generative foundation models where the output, despite being grammatically correct, lacks diversity and originality, frequently repeating phrases, sentence structures, or ideas.
- Often due to sampling strategies (e.g., greedy decoding).
- Can be mitigated by using diverse decoding methods like nucleus sampling or beam search with penalty.
- Reduces perceived creativity and usefulness of the generated content.
Memory trick: Repetitive generation lacks true creativity.
On-premises FM Deployment for Data Sovereignty
Flip cardThe practice of hosting and operating foundation models within an organization's own physical or private cloud infrastructure to meet strict data residency, privacy, and security mandates.
- Ensures data never leaves the organization's control.
- Requires significant hardware, infrastructure, and operational expertise.
- Often chosen by organizations in highly regulated industries (e.g., finance, healthcare, legal).
Memory trick: On-premises: Your data, your fortress, your rules.
AWS Real-Time Spoken Language Translation
Flip cardAchieving real-time translation of spoken audio into text in multiple target languages by combining Amazon Transcribe and Amazon Translate.
- Amazon Transcribe converts live audio to text.
- Amazon Translate translates the text into desired languages.
- Enables global audiences to consume spoken content in their native language.
Memory trick: Transcribe the speech, then Translate the text for global reach.
Amazon Kinesis
Flip cardA suite of services for collecting, processing, and analyzing streaming data in real time, enabling immediate insights and rapid responses.
- Handles massive data streams from various sources.
- Low-latency processing for real-time analytics.
- Includes Kinesis Data Streams, Firehose, Video Streams, and Data Analytics.
Memory trick: Kinesis keeps the stream flowing for real-time insights.