CompTIA Data+ (DA0-002) flashcards
149 free flashcards. Tap a card to flip it.
Retention Schedule
Flip cardA governance policy specifying how long each category of records must be retained before secure disposal, based on legal, regulatory, or business needs.
- Different record types may have different mandated periods (e.g., 7 years for tax records)
- After the period expires, records are typically purged unless a legal hold applies
- Helps reduce storage costs and legal exposure from over-retention
Memory trick: Retention schedule = a countdown timer โณ on a filing cabinet ๐๏ธ
SQL Case-Insensitive Search
Flip cardPerforming a search in SQL that ignores the case of characters, often achieved by converting strings to a common case (e.g., lowercase) before comparison.
- Typically uses `LOWER()` or `UPPER()` functions.
- Some SQL dialects offer specific operators like `ILIKE` (PostgreSQL).
- Crucial for robust data matching and filtering.
Memory trick: Filter carefully, get precise results.
SQL CASE Statement
Flip cardA conditional expression in SQL that allows you to define different results based on various conditions, similar to 'if-then-else' logic in programming.
- Used for conditional logic in SELECT, WHERE, ORDER BY, and GROUP BY clauses.
- Can map multiple input values to a single output value.
- Essential for data transformation, categorization, and handling complex business rules.
Memory trick: SQL conditions are like a CASE: when this, then that, else something else.
Kruskal-Wallis H-test
Flip cardA non-parametric test used to compare the distributions of a continuous or ordinal variable for two or more independent groups. It is an alternative to One-Way ANOVA when assumptions are violated.
- Non-parametric (does not assume normality).
- Compares medians or ranks, not means.
- Suitable for ordinal data or non-normal interval/ratio data.
- Extension of the Mann-Whitney U test for more than two groups.
Memory trick: Kruskal 'K'eeps it non-parametric.
Data Normalization (Cleansing)
Flip cardThe process of organizing data in a database to eliminate redundancy and improve data integrity, often involving standardizing inconsistent values to a common format.
- Aims to reduce data redundancy.
- Ensures data consistency and integrity.
- Involves standardizing formats, units, and representations.
Memory trick: Transforming data makes it shine.
Data Governance Council
Flip cardA cross-functional oversight body that sets enterprise data policy, resolves ownership disputes, and prioritizes governance initiatives.
- Includes IT, legal, compliance, and business stakeholders
- Sets strategic direction; stewards execute day-to-day rules
- Resolves cross-department disagreements over data ownership
Memory trick: Council sits at the top like a courtroom judge deciding data disputes
Decision Tree Analysis
Flip cardA supervised learning method used for classification and regression that builds a model in the form of a tree structure, where internal nodes represent tests on an attribute, branches represent outcomes of the test, and leaf nodes represent class labels or predicted values.
- Visually interpretable for explaining predictive factors.
- Handles both numerical and categorical data.
- Identifies the most influential variables and their thresholds.
Memory trick: A Decision Tree helps managers see which branches lead to churn, making it easy to decide what to do.
Mode
Flip cardThe value that appears most frequently in a dataset.
- Can be used for all types of data (nominal, ordinal, interval, ratio).
- A dataset can have one mode (unimodal), multiple modes (multimodal), or no mode.
- Not affected by extreme values (outliers).
Memory trick: Mode: The MOST popular number in the crowd.
Ordinal Data
Flip cardA type of categorical data where the categories have a natural, meaningful order or ranking, but the differences between categories are not uniform or quantifiable.
- Categories have a clear order
- Intervals between categories are not equal or precisely measurable
- Cannot perform arithmetic operations like addition or averaging
Memory trick: N-O-I-R: N-o Order, O-rder, I-ntervals, R-atio.
Completeness (Data Quality Dimension)
Flip cardThe degree to which all required data values are present in a dataset, with no missing mandatory fields.
- Measured as % of populated required fields vs total expected
- Missing/null values in mandatory fields reduce completeness
- Different from accuracy, which measures correctness of present values
Memory trick: 'CACTUS' - Completeness, Accuracy, Consistency, Timeliness, Uniqueness, Validity Standards
Seasonality Analysis
Flip cardThe process of identifying and quantifying recurring patterns or cycles in time series data that repeat over a fixed period.
- Part of time series decomposition.
- Patterns repeat at regular intervals (e.g., daily, weekly, monthly).
- Important for accurate forecasting and understanding underlying dynamics.
Memory trick: TIME Series: Trend, Season, Cycle, Irregular.
ANOVA (Analysis of Variance)
Flip cardA statistical test used to compare the means of three or more independent groups to determine if there is a statistically significant difference between them.
- Assumes normality and homogeneity of variances (equal variances across groups).
- Uses the F-statistic to test the null hypothesis that all group means are equal.
- If significant, further post-hoc tests are needed to identify which specific groups differ.
Memory trick: ANOVA is like a 'committee meeting' for means, checking if all groups agree or if someone stands out.
Attribute-Based Access Control (ABAC)
Flip cardAn access control model that grants permissions based on evaluating multiple attributes of the user, resource, and environment (e.g., department, time, location) against policy rules.
- More granular and dynamic than RBAC
- Can incorporate context such as time of day or location
- Commonly used where access needs change frequently, like healthcare
Memory trick: ABAC = Attributes Assembled Before Access Control decision
Interactive Dashboard
Flip cardA visual display of data that allows users to manipulate parameters, filter data, and drill down into details to explore information relevant to their specific questions.
- Empowers self-service data exploration.
- Accommodates diverse user needs.
- Enhances engagement and deeper insights.
Memory trick: Many eyes, many needs, one report that truly feeds.
ELT Load Phase (Minimal Transformation)
Flip cardIn an ELT (Extract, Load, Transform) pipeline, the Load phase typically involves moving raw data directly into a target system (like a data lake) with minimal, often non-destructive, transformations to ensure immediate availability and preserve raw data integrity.
- Prioritizes speed of data ingestion and availability of raw data.
- Transformations are usually limited to schema enforcement, data type conversion, and basic cleansing.
- Complex transformations are deferred to the 'Transform' stage, usually performed within the target data lake or warehouse.
Memory trick: ELT is like a truck: Extract, Load quickly to the lake, then Transform as needed.
SQL Date String Conversion
Flip cardThe process of converting a string representation of a date into a proper date data type within SQL, often requiring a format specification.
- Essential for consistent date handling and calculations.
- Functions like `TO_DATE()` or `STR_TO_DATE()` specify input format masks.
- Different SQL dialects have specific functions for this task.
Memory trick: Converting strings to dates needs a clear map.
F1-Score
Flip cardThe harmonic mean of precision and recall, used as a single metric to evaluate binary classification models, especially with imbalanced datasets.
- Balances false positives (precision) and false negatives (recall).
- Ranges from 0 to 1, with 1 being perfect.
- Useful when both types of errors are costly.
Memory trick: F1 balances the 'F'ine line of errors.
Recall (Sensitivity)
Flip cardThe proportion of actual positive instances that were correctly identified by the model. It focuses on minimizing false negatives.
- Formula: True Positives / (True Positives + False Negatives).
- Important when the cost of a false negative is high (e.g., fraud detection, medical diagnosis).
- Also known as the True Positive Rate (TPR).
Memory trick: Recall 'R'emembers all the real ones.
Ratio Data
Flip cardA type of quantitative data that has a true zero point, meaning zero indicates the absence of the measured quantity, and ratios between values are meaningful.
- Has a true zero point.
- Ratios between values are meaningful.
- Allows for all arithmetic operations (addition, subtraction, multiplication, division).
- Includes measurements like height, weight, age, and count.
Memory trick: Imagine a 'Ratio' as a 'ruler' that starts at zero and lets you compare sizes.
SQL Data Transformation
Flip cardSQL data transformation involves applying functions and operations to modify data from its raw form into a desired, standardized, or optimized format for analysis or storage.
- Includes string manipulation, date formatting, type casting, and conditional logic.
- Crucial for data quality and consistency.
- Often performed during ETL/ELT processes.
Memory trick: Case first, then fix specific patterns!
Static Data Masking
Flip cardA technique that permanently replaces sensitive data with realistic but fictitious values in a copy of a dataset, typically for use in non-production environments like testing or development.
- Creates a separate masked copy, unlike dynamic masking of live data
- Preserves data format/referential integrity for testing
- No original sensitive value remains recoverable in the masked copy
Memory trick: Dynamic masks live at query time; Static masks a frozen copy forever
Least Privilege
Flip cardA security principle stating users should be granted only the minimum access rights needed to perform their job functions.
- Reduces attack surface and insider risk
- Additional access requires approval
- Complements RBAC and ABAC as an overarching goal
Memory trick: LSND: Least privilege, Separation of duties, Need-to-know, Deny by Default
SQL Regular Expressions
Flip cardSQL regular expressions (REGEXP_LIKE, RLIKE, REGEXP_MATCH, etc.) provide advanced pattern matching capabilities for searching and manipulating strings based on complex patterns.
- More powerful and flexible than LIKE for complex patterns.
- Uses standard regular expression syntax.
- Essential for robust data validation and parsing.
Memory trick: Simple LIKE, Complex REGEX!
Data Lineage
Flip cardDocumentation that tracks the flow of data from its original source through every transformation and system to its final destination, supporting root-cause analysis and impact assessment.
- Traces data from source to report/consumption
- Essential for troubleshooting and auditing
- Part of broader metadata management
Memory trick: Lineage is the 'family tree' tracing data's journey.
SQL Join Optimization
Flip cardTechniques used to improve the performance of SQL queries involving JOIN operations, especially on large datasets.
- Filtering data before joining reduces the dataset size.
- Proper indexing on join columns is crucial.
- Using appropriate join types can impact performance.
- Data type consistency prevents implicit conversions.
Memory trick: Optimizing queries speeds up data insights.
Geographic Heatmap
Flip cardA visualization that displays the density or magnitude of data points across a geographical area, typically using color intensity to represent values.
- Shows spatial distribution of data.
- Uses color variation to represent data density or value.
- Ideal for identifying hot spots or clusters on a map.
- Can visualize various metrics like population density, sales concentration, or event frequency.
Memory trick: Heatmaps show HOT spots on the MAP!
F1-Score for Imbalanced Data
Flip cardThe F1-Score is the harmonic mean of precision and recall, specifically useful for evaluating classification models on datasets where class distribution is imbalanced.
- Calculated as 2 * (Precision * Recall) / (Precision + Recall).
- Provides a single score balancing true positives, false positives, and false negatives.
- Penalizes models that perform poorly on either precision or recall.
Memory trick: For imbalanced data, F1-Score is like a tightrope walker, balancing Precision and Recall perfectly.
Unstructured Data
Flip cardInformation that does not have a predefined data model or is not organized in a pre-defined manner. It typically consists of text and multimedia content.
- No fixed schema or format.
- Examples: text documents, emails, social media posts, audio, video.
- Requires advanced techniques (NLP, machine learning) for analysis.
- Accounts for the vast majority of enterprise data.
Memory trick: Structured is neat, Semi is tagged, Unstructured is free.
Paired Samples t-test
Flip cardA statistical test used to determine if there is a significant difference between the means of two related groups or samples. It is typically applied when the same subjects are measured twice (e.g., before and after an intervention).
- Compares means of two related samples.
- Accounts for dependency between measurements.
- Requires continuous data for the dependent variable.
Memory trick: Same subjects, paired comparison, that's the t-test you've aired!