CompTIA Data+ (DA0-002) flashcards
149 free flashcards. Tap a card to flip it.
ELT Paradigm
Flip cardA data integration approach where data is first extracted from sources, loaded directly into a target data store (like a data lake), and then transformed within that same data store.
- Optimized for cloud and big data environments.
- Allows for faster initial data loading.
- Leverages the scalability and processing power of the target system.
Memory trick: Integration choices shape data flow.
Pseudonymization
Flip cardA reversible de-identification technique that replaces identifying data fields with artificial identifiers (pseudonyms), while a separately secured mapping table allows authorized re-identification.
- Reversible, unlike anonymization
- Mapping table must be securely stored separately
- Still considered personal data under regulations like GDPR
Memory trick: Pseudonymization wears a 'mask' that can be removed with the right key.
Consistency and Alignment (Design)
Flip cardThe principle of maintaining uniform visual elements (layout, colors, fonts, scales) across a dashboard or report to improve readability, comparability, and user experience.
- Reduces cognitive load and confusion.
- Facilitates quick comparison and trend identification.
- Creates a cohesive and professional appearance.
Memory trick: Trends and quirks, clearly seen, with consistent design, sharp and keen.
Data Sovereignty
Flip cardThe principle that data is subject to the laws and governance structures of the country or jurisdiction in which it is collected or stored.
- Often drives requirements for data residency/localization
- Common in GDPR and similar cross-border data laws
- Distinct from data governance, which is an internal management framework
Memory trick: Sovereignty=law of the land, Residency=where it lives, Localization=must stay local
Mean Imputation
Flip cardA data cleansing technique where missing numerical values in a dataset are replaced with the average (mean) of the available values for that specific feature.
- Preserves data volume by not deleting rows.
- Introduces less bias than deleting rows if data is missing at random.
- Can reduce variance and distort relationships if used excessively or on non-random missing data.
Memory trick: Missing pieces need a smart fill, not a complete spill.
JSON (JavaScript Object Notation)
Flip cardA lightweight, human-readable data interchange format that uses human-readable text to transmit data objects consisting of attribute–value pairs and array data types.
- Lightweight and commonly used for web applications.
- Supports nested structures and arrays.
- Human-readable and easy to parse.
- Language-independent, widely supported across programming languages.
Memory trick: JSON is 'J'ust 'S'imple 'O'bject 'N'otation.
Treemap
Flip cardA visualization that displays hierarchical data as a set of nested rectangles, where the size of each rectangle is proportional to its value, and colors can represent categories or other attributes.
- Effective for showing parts of a whole with many categories.
- Good for identifying dominant categories at a glance.
- Can represent hierarchical structures efficiently.
Memory trick: Many pieces, one big whole, treemap helps to meet your goal.
Problem-Solution Framework
Flip cardA storytelling structure that identifies a problem, proposes a solution, and then presents evidence or results demonstrating the solution's effectiveness.
- Highly effective for business presentations.
- Engages the audience by addressing their pain points.
- Clearly demonstrates the value and impact of the analysis.
Memory trick: Problems Solved, Stories Evolved.
Gauge Chart
Flip cardA circular or semi-circular chart that displays a single data point's value against a defined range or target, often using color to indicate performance status.
- Ideal for displaying KPI performance against targets.
- Provides immediate visual feedback on status.
- Effective for dashboards requiring 'at-a-glance' insights.
Memory trick: Target met, or far astray? The gauge will show you, come what may.
Pie Chart
Flip cardA circular statistical graphic divided into slices to illustrate numerical proportion. Each slice represents a category's share of the total.
- Best for showing parts of a whole.
- Effective with a small number of categories (typically 2-6).
- Difficult to compare precise values between slices.
Memory trick: Pie slices show parts, like pizza parts.
Data Sensitivity
Flip cardThe degree to which data requires protection from unauthorized access, disclosure, alteration, or destruction.
- Highly sensitive data often includes PII, financial records, health information, and intellectual property.
- Governs the level of security controls, access restrictions, and compliance requirements.
- Mismanagement can lead to legal penalties, reputational damage, and financial loss.
Memory trick: Data is defined by its 'V's and its 'S' – how much, how fast, what kind, and how secret.
AI Fairness
Flip cardThe principle that AI systems should produce unbiased and equitable outcomes for all individuals and demographic groups, avoiding discrimination.
- Involves identifying and mitigating biases in training data and model algorithms.
- Crucial for ethical AI development, especially in sensitive applications like lending or hiring.
- Requires careful consideration of metrics beyond just accuracy, such as disparate impact.
Memory trick: Ethical AI is like building a fair and transparent judge, not a biased black box.
SQL WHERE Clause
Flip cardThe WHERE clause in SQL is used to extract only those records that fulfill a specified condition or conditions.
- Filters individual rows before aggregation.
- Can combine multiple conditions using AND, OR, NOT operators.
- Commonly used with SELECT, UPDATE, DELETE statements.
Memory trick: Where to filter? Before the group, at the row!
Timeliness (Data Quality Dimension)
Flip cardThe degree to which data is up to date and available within the timeframe required for its intended use.
- Data can be accurate and complete yet still fail timeliness if it's stale
- Common issue in real-time dashboards fed by batch updates
- Improved via streaming/real-time data pipelines instead of nightly batches
Memory trick: 'CACTUS' - Completeness, Accuracy, Consistency, Timeliness, Uniqueness, Validity Standards
Time-Series Database (TSDB)
Flip cardA database optimized for storing and retrieving time-stamped data points, ideal for metrics, events, and sensor data.
- Optimized for time-indexed data
- High write throughput for sequential data
- Efficient time-range queries and aggregations
- Often includes data retention policies
Memory trick: Time-series is the T-rue S-olution for T-ime-stamped S-ensors.
Geographic Bubble Map (with Proportions)
Flip cardA map visualization where bubbles are placed on geographical locations, with bubble size representing a quantitative value and often internal segments or colors within the bubble representing proportional contributions.
- Combines geographical context with quantitative data.
- Bubble size often represents magnitude (e.g., total sales).
- Can incorporate secondary proportional data (e.g., market share by product) within each bubble.
Memory trick: Bubbles on a map, telling a proportional story.
KPI Tile (Key Performance Indicator Tile)
Flip cardA dashboard element designed to prominently display a single, crucial metric, often accompanied by its value, target, and status (e.g., trend or conditional color).
- Focuses on a single, important metric.
- Often includes a target or benchmark.
- Can use conditional formatting to indicate status (e.g., good/bad).
Memory trick: KPI Tiles are Key to Performance Indicators, Clearly!
Pearson Correlation Coefficient (r)
Flip cardA measure of the linear correlation between two sets of data.
- Ranges from -1 to +1.
- Sign indicates direction (positive/negative).
- Magnitude indicates strength (closer to ±1 is stronger).
- Does not imply causation.
Memory trick: CORRELATION: Does it go together, and how TIGHTLY?
Right to Erasure (GDPR)
Flip cardA GDPR data subject right, also called the 'right to be forgotten,' allowing individuals to request deletion of their personal data when it is no longer needed or other legal conditions apply.
- One of eight GDPR data subject rights
- Not absolute — exceptions include legal obligations
- Requires organizations to delete data from all relevant systems
Memory trick: 'Erase' = the 'forget me' button under GDPR.
P-value and Significance Level
Flip cardThe p-value is the probability of observing data as extreme as, or more extreme than, the observed data, assuming the null hypothesis is true. The significance level (alpha) is the threshold below which the null hypothesis is rejected.
- If p-value < alpha, reject the null hypothesis.
- If p-value ≥ alpha, fail to reject the null hypothesis.
- Alpha is typically set at 0.01, 0.05, or 0.10.
Memory trick: P-value low, null must go! P-value high, null can fly!
ELT (Extract, Load, Transform)
Flip cardA data integration approach where data is first extracted from sources, loaded into a target system (like a data lake), and then transformed within that system.
- Favored for big data and cloud environments.
- Allows for faster data loading.
- Leverages the processing power of the target system for transformations.
Memory trick: Fast data flow needs smart processing.
Data Ink Ratio
Flip cardA design principle by Edward Tufte stating that the proportion of a graphic's ink devoted to the non-redundant display of data-information should be maximized.
- Maximize 'data-ink': ink used to display data.
- Minimize 'non-data-ink': ink used for labels, borders, frames, etc.
- Aims for clarity and efficiency in data visualization.
Memory trick: Data Ink: Maximize Data, Minimize Clutter!
Multi-Line Chart
Flip cardA line chart that displays multiple distinct data series, each represented by its own line, on the same set of axes, typically used for comparing trends over time.
- Compares trends of multiple categories over time.
- Each category has its own line.
- Effective for identifying divergences or convergences in trends.
- Requires a continuous variable (time) on the x-axis.
Memory trick: Many Lines, Many Trends, Easy to See What's Sending!
Data Disposal (Destruction)
Flip cardThe final stage of the data lifecycle in which data is permanently and irreversibly destroyed after its retention period expires, using methods like shredding or cryptographic erasure.
- Occurs only after retention requirements/legal holds are satisfied
- Methods include physical destruction and cryptographic erasure
- Failure to properly dispose of data can create compliance and security risk
Memory trick: Create, Store, Use, Share, Archive, then finally Destroy for good
Visual Cues (Conditional Formatting)
Flip cardThe use of visual elements like color, icons, size, or shape to highlight specific data points or conditions, drawing the viewer's attention to important information or trends.
- Provides immediate, intuitive understanding.
- Reduces cognitive load by eliminating need to read exact numbers.
- Commonly used for alerts, status indicators, and comparison.
- Examples include color scales, traffic light indicators, and arrow icons.
Memory trick: Visual Cues give INSTANT CLUES!
Key-Value Store
Flip cardA simple NoSQL database that stores data as a collection of key-value pairs, where each key is unique and used to retrieve its associated value. It's optimized for high-speed read/write operations.
- Simplest NoSQL data model.
- Data stored as unique key-value pairs.
- Extremely fast for read and write operations by key.
- Highly scalable, often used for caching, session management, and real-time data.
Memory trick: A 'Key-Value' store is like a simple 'Key' to unlock a 'Value' quickly.
ETL Transformation
Flip cardThe phase in an ETL process where raw data is converted, cleansed, and aggregated into the desired format for analysis and storage.
- Occurs between extraction and loading.
- Ensures data quality and consistency.
- Involves data cleansing, standardization, and aggregation.
Memory trick: Every Tidy Loader ensures data is ready.
SQL String Cleansing
Flip cardThe process of using SQL string functions to clean and standardize text data by removing unwanted characters, spaces, or inconsistencies.
- Common tasks include removing leading/trailing spaces, extra internal spaces, special characters, and converting case.
- Functions like TRIM, LTRIM, RTRIM, REPLACE, and sometimes regular expressions are frequently used.
- Ensures text data is consistent and ready for analysis or storage.
Memory trick: Cleaning text with SQL: TRIM the ends, REPLACE the messy middle.
Empirical Rule (68-95-99.7 Rule)
Flip cardA rule stating that for a normal distribution, nearly all data falls within three standard deviations of the mean.
- Approx. 68% of data within 1 standard deviation (μ ± 1σ).
- Approx. 95% of data within 2 standard deviations (μ ± 2σ).
- Approx. 99.7% of data within 3 standard deviations (μ ± 3σ).
Memory trick: NORMAL bell, 68-95-99.7 is the RULE!
Data Mapping
Flip cardThe process of creating a correspondence between data elements from a source system to target data elements in another system, often involving conversion or transformation rules.
- Essential for data integration and migration.
- Defines how data fields relate and convert.
- Can include value-level transformations and data type conversions.
Memory trick: Transform data: shape it, combine it, or make it fit just right.
Pearson Correlation Coefficient
Flip cardA statistical measure that quantifies the strength and direction of a linear relationship between two continuous variables, ranging from -1 to +1.
- Values range from -1 (perfect negative linear correlation) to +1 (perfect positive linear correlation).
- A value of 0 indicates no linear correlation.
- Only measures linear relationships; non-linear relationships may exist even with r=0.
Memory trick: Pearson Correlation is like seeing if two friends always walk in the same direction and how close they stay.
Standard Deviation
Flip cardA measure of the amount of variation or dispersion of a set of values.
- Indicates how spread out numbers are from the average (mean).
- A low standard deviation means values are close to the mean.
- A high standard deviation means values are spread out over a wider range.
Memory trick: STD DEV: How SPREAD OUT your data is.
Seasonality (Time Series)
Flip cardA component of time series data that describes regular and predictable patterns or fluctuations that recur over a fixed period, such as daily, weekly, monthly, or annually.
- Predictable and recurring patterns.
- Occurs within a fixed period (e.g., weekly, quarterly).
- Can be removed or modeled for better forecasting.
Memory trick: Seasonality is like the seasons themselves – always coming back at predictable times.
Robustness of Median
Flip cardThe median's ability to remain largely unchanged by extreme values (outliers) in a dataset, making it a robust measure of central tendency for skewed distributions.
- Unaffected by magnitude of outliers, only their position.
- Preferred for income, property values, or other skewed distributions.
- Contrasts with the mean, which is pulled by outliers.
Memory trick: Median for money, because it's not swayed by the rich.
Data Standardization
Flip cardThe process of transforming data into a common format, scale, or structure to ensure consistency and comparability across different sources or within a single dataset.
- Essential for integrating data from multiple sources.
- Includes converting units, data types, and formatting conventions.
- Often combined with data validation to enforce business rules.
Memory trick: Clean data is like a polished gem, ready for its big show.
K-Means Clustering
Flip cardAn unsupervised machine learning algorithm that partitions a dataset into 'k' distinct, non-overlapping subgroups (clusters), where each data point belongs to the cluster with the closest mean (centroid).
- Unsupervised: no labeled target variable required.
- Aims to minimize intra-cluster variance and maximize inter-cluster variance.
- Requires specifying the number of clusters (k) beforehand.
Memory trick: K-Means is like telling a group of kids to sort themselves into K teams based on how similar they feel.
Accuracy Rate Calculation
Flip cardA metric expressing the percentage of records that correctly match a trusted reference source, calculated as (matching records ÷ total records) × 100.
- Accuracy differs from validity (accuracy checks truth, not just format)
- Formula: matching records ÷ total records × 100
- Requires a trusted reference source for comparison
Memory trick: Divide the matches by the total, then dress it up in a percent hat.
Graph Database
Flip cardA type of NoSQL database that uses graph structures (nodes, edges, and properties) to store and represent highly interconnected data, making it ideal for managing and querying relationships.
- Stores data as nodes (entities) and edges (relationships).
- Optimized for complex relationship queries.
- Ideal for social networks, recommendation engines, fraud detection.
- Can traverse many-to-many relationships efficiently.
Memory trick: Think 'Graph' for 'G'reat 'R'elationship 'A'nalysis.
Text Normalization
Flip cardA crucial step in natural language processing (NLP) that transforms raw text into a standardized and clean format, making it more suitable for analysis. It includes tasks like lowercasing, removing punctuation, special characters, and standardizing informal language.
- Aims to reduce variability in text data.
- Essential for tasks like sentiment analysis, topic modeling, and text classification.
- Often involves regular expressions for pattern-based cleaning.
Memory trick: Cleaning text for NLP is like polishing a gem: normalization makes it shine.
Deduplication
Flip cardA data cleansing process that identifies and removes duplicate records or entries from a dataset, ensuring uniqueness based on specific keys or attributes.
- Crucial for maintaining data integrity and accuracy.
- Can be performed based on exact matches or fuzzy matching algorithms.
- Prevents overcounting and skewed analysis in reporting and models.
Memory trick: Clean data, clear insights: no mess, no fuss, just truth.
Outlier Resistance (Central Tendency)
Flip cardThe ability of a measure of central tendency to remain stable and representative of the data's center despite the presence of extreme values (outliers). The median is highly resistant.
- Median: Most resistant to outliers.
- Mean: Least resistant, heavily influenced by outliers.
- Mode: Generally resistant, but may not be unique or central for all data types.
Memory trick: Median's 'M'iddle position means it's 'M'ost resistant.
NoSQL Database
Flip cardA non-relational database that provides a mechanism for storage and retrieval of data that is modeled in means other than the tabular relations used in relational databases.
- Handles schema-less data
- High scalability and flexibility
- Suitable for large volumes of rapidly changing data
Memory trick: NoSQL is N-o-t just SQL, it's N-ew O-ptions for S-calability and Q-uick L-oads.
SQL INNER JOIN
Flip cardAn INNER JOIN returns only the rows that have matching values in both tables involved in the join. It effectively selects the intersection of the two tables based on the join condition.
- Most common type of join.
- Requires a matching key in both tables.
- Excludes rows that do not have a match in the other table.
Memory trick: Inner Join: Only the overlap, no stragglers!
Median
Flip cardThe middle value in a dataset that is ordered from least to greatest.
- Robust to outliers and skewed distributions.
- For an odd number of values, it's the single middle value.
- For an even number of values, it's the average of the two middle values.
Memory trick: MEDIAN: The middle path, ordered first!
Separation of Duties (SoD)
Flip cardA control principle that divides critical tasks, such as authorization, execution, and review, among multiple individuals to reduce the risk of fraud or error.
- Prevents any single person from controlling an entire process
- Commonly applied to financial transactions and approvals
- Complements least privilege but focuses on task division, not access scope
Memory trick: No one person should both write the check and sign it.
Consistent Visual Language
Flip cardThe use of uniform design elements (colors, fonts, icons, chart types) and explicit explanations (annotations) across a visualization or report to aid comprehension and reduce cognitive load.
- Enhances clarity and understanding.
- Builds familiarity and reduces learning curve.
- Crucial for communicating technical data to diverse audiences.
Memory trick: Speak two tongues, with visuals bright, make data's meaning clear as light.
Predictive Modeling
Flip cardA statistical or machine learning technique used to forecast future outcomes or probabilities based on historical data.
- Identifies patterns in past data to predict future behavior.
- Common applications include fraud detection, risk assessment, and customer churn prediction.
- Often involves classification or regression algorithms.
Memory trick: AI tasks are like different superpowers: some let you see, some let you speak, and some let you predict the future.
Hero Metric
Flip cardA single, most important key performance indicator (KPI) that clearly communicates the primary success or failure of an initiative.
- Simplifies complex outcomes.
- Provides a clear focus for storytelling.
- Easily understood by diverse audiences.
Memory trick: Data's tale, when told with grace, leaves a lasting, clear embrace.
Bar Chart with Annotations
Flip cardA chart type that uses rectangular bars to represent discrete categories, with the length of each bar proportional to the value it represents, often enhanced with text labels for additional context like percentages.
- Excellent for comparing individual category values.
- Easy to rank and identify top/bottom performers.
- Annotations can add proportional context without needing a separate chart.
Memory trick: Bars Labeled, Performance Unraveled.
Role-Based Access Control (RBAC)
Flip cardAn access control model that assigns permissions to roles corresponding to job functions, and users inherit permissions by being assigned to a role.
- Simplifies management by grouping permissions into roles
- New users gain access instantly upon role assignment
- Contrasts with ABAC, which uses multiple contextual attributes
Memory trick: DAC=discretion, MAC=military, RBAC=role, ABAC=attribute soup.
Mann-Whitney U Test
Flip cardA non-parametric statistical hypothesis test used to compare the distributions of two independent samples to determine if they are significantly different.
- Used for ordinal or non-normally distributed interval/ratio data.
- Compares two independent groups.
- A non-parametric alternative to the independent samples t-test.
Memory trick: Mann-Whitney helps when data is not normal, like two different teams playing at different skill levels.
Anonymization
Flip cardThe irreversible process of removing or altering personal identifiers so that individuals cannot be re-identified from the data, even by the data holder.
- Irreversible, unlike pseudonymization or tokenization
- Often involves aggregation or generalization of fields
- Fully anonymized data typically falls outside privacy regulation scope
Memory trick: Mask hides, Pseudonym swaps (reversible), Anonymize erases forever, Token vaults, Aggregate blends
Confidence Interval
Flip cardA range of values, derived from sample statistics, that is likely to contain the value of an unknown population parameter.
- Expressed with a confidence level (e.g., 90%, 95%, 99%).
- Indicates the reliability of an estimate.
- Wider intervals indicate greater uncertainty.
Memory trick: CI: How CONFIDENT are we that our NET catches the true mean?
Median for Skewed Data
Flip cardThe median is the middle value in a sorted dataset, often preferred over the mean for skewed distributions as it is less affected by outliers.
- Divides data into two equal halves.
- Resistant to extreme values.
- Best for ordinal or interval data with skew.
Memory trick: Mean is for averages, Median is for middle, Mode is for most.
SQL Data Validation
Flip cardSQL data validation involves using SQL queries to check data against predefined rules, constraints, or patterns to ensure its accuracy, consistency, and integrity.
- Uses WHERE clauses, LIKE/REGEXP patterns, IS NULL, and comparison operators.
- Essential for data quality and reliability.
- Can be used to identify records needing cleansing or correction.
Memory trick: Null, Empty, Specific Bad, or Bad Pattern!
Mean (Arithmetic Mean)
Flip cardThe average of a set of numbers, calculated by summing all values and dividing by the count of values.
- Sensitive to outliers and skewed distributions.
- Best for symmetrically distributed data.
- Represents the 'balance point' of the data.
Memory trick: Mean, Median, Mode: What's the Middle Story?
Significance Level (α)
Flip cardThe probability of rejecting the null hypothesis when it is actually true (Type I error). It is set by the researcher before hypothesis testing.
- Also known as alpha (α).
- Commonly set at 0.05 or 0.01.
- Determines the critical region for hypothesis tests.
Memory trick: Alpha sets the 'A'-cceptable error rate.
Discretionary Access Control (DAC)
Flip cardAn access control model in which the owner of a resource determines who is granted access and what permissions they receive.
- Owner controls permissions, not a central authority
- Common in file systems (e.g., Windows NTFS permissions)
- Contrasts with MAC, which enforces centrally defined policies
Memory trick: DAC=owner decides, MAC=military/central rules, RBAC=role badge, ABAC=attribute checklist
Data Smoothing
Flip cardTechniques used to remove noise or short-term fluctuations from data, especially time series, to reveal underlying trends or patterns more clearly.
- Common methods: moving averages, exponential smoothing.
- Reduces volatility and highlights long-term trends.
- Helps in forecasting and understanding overall patterns.
Memory trick: Smooth the bumps to see the path.
Independent Samples t-test
Flip cardA statistical test used to determine if there is a significant difference between the means of two independent groups.
- Compares means of two separate populations/groups.
- Dependent variable is continuous, independent variable is categorical (two levels).
- Assumes independence of observations and often normality and equal variances.
Memory trick: T-TEST: Are THESE TWO groups different?