Step2Study
IT & TechnologyDA0-002100% Free

CompTIA Data+ (DA0-002)

Practice bank
230 Qs
Real exam
90 Qs
Time limit
90 min
Passing
675 on a 100–900 scale

Exam blueprint

Data Concepts and Environments
15%
Data Mining
25%
Data Analysis
23%
Visualization
23%
Data Governance, Quality and Controls
14%

Practice

Untimed · instant feedback · 4 practice tests of 90 questions

Questions per test

Custom practice

Flashcard on every question Mental map when you miss

Exam simulation

4 timed tests · 90 questions each · 90 min · pass 75% · 230 questions in the bank

+50 XP per test · +100 XP for a pass

Random simulation (weighted by domain)

Everything is open to everyone. Create a free account to save scores, XP, badges and get progress emails.

Free study resources

All resources →

Part of a learning path

Study with friends

Challenge a friend to beat your score.

CompTIA Data+ (DA0-002) practice test questions

Sample questions from the 230-question bank, with answers and explanations.

All questions
  1. 1. A business analyst is preparing a report to compare the average sales performance of three distinct product categories across four different regions. The goal is to easily identify which product category performs best in each region and to quickly spot any underperforming categories. Which type of chart would be MOST effective for this visualization?

    Visualization

    • A. Line Chart
    • B. Grouped Bar Chart
    • C. Scatter Plot
    • D. Pie Chart
    Show answer

    B. Grouped Bar Chart

    A grouped bar chart is ideal for comparing multiple categories (product categories) across different groups (regions), allowing for easy visual comparison of performance within each region.

  2. 2. A data analyst is performing an exploratory data analysis on a sales dataset. They notice that the 'Revenue' column contains a few values that are significantly higher than the vast majority of other entries, potentially skewing statistical measures like the mean and standard deviation. These extreme values, while possibly legitimate, are distorting the overall distribution. Which data cleansing technique is most appropriate to mitigate the impact of these extreme values on statistical analysis without necessarily deleting them?

    Data Mining

    • A. Data normalization or winsorization
    • B. Schema validation
    • C. Deduplication
    • D. Missing value imputation
    Show answer

    A. Data normalization or winsorization

    Data normalization (like log transformation) can reduce the impact of skewed distributions and outliers by transforming the data scale. Winsorization, specifically, caps extreme values at a certain percentile, effectively reducing their influence without removal. Both are appropriate for mitigating the impact of outliers while retaining data.

  3. 3. A business intelligence team needs to analyze historical sales data, customer demographics, and product inventory to identify trends and create reports for strategic decision-making. The data comes from various operational systems and needs to be cleaned, transformed, and loaded into a central repository optimized for complex analytical queries. Which data environment is best suited for this purpose?

    Data Concepts and Environments

    • A. Data Warehouse
    • B. Data Lake
    • C. NoSQL Database
    • D. Operational Database
    Show answer

    A. Data Warehouse

    A data warehouse is specifically designed to store integrated, historical data from various sources, optimized for analytical querying and reporting, making it ideal for business intelligence and strategic decision-making.

  4. 4. A company's production database must allow full, unmasked access to Social Security numbers for the compliance team, while call center representatives querying that same live database see only the last four digits of each SSN at the moment of query, with no change made to the actual stored values. Which technique satisfies this requirement?

    Data Governance, Quality and Controls

    • A. Pseudonymization
    • B. Dynamic data masking
    • C. Static data masking
    • D. Encryption at rest
    Show answer

    B. Dynamic data masking

    Dynamic data masking obscures sensitive data in real time based on the requesting user's role or privileges, without altering the underlying stored data, so different users can see different views of the same live record. Static masking permanently alters copies of data (e.g., for test environments), pseudonymization permanently replaces identifiers in a dataset, and encryption at rest protects data only while stored, not selectively by viewer.

  5. 5. A data analyst is designing a dashboard that displays real-time customer support metrics, such as average response time and resolution rate. The dashboard needs to be constantly monitored by team leads to identify immediate issues. Which design principle is MOST critical to ensure the dashboard remains effective and actionable in this dynamic environment?

    Visualization

    • A. Requiring manual data refresh for accuracy.
    • B. Incorporating a highly detailed data table for every metric.
    • C. Extensive use of static images and complex infographics.
    • D. Utilizing clear, intuitive visual cues, such as color-coding for thresholds.
    Show answer

    D. Utilizing clear, intuitive visual cues, such as color-coding for thresholds.

    For a real-time dashboard requiring immediate issue identification, clear and intuitive visual cues like color-coding for thresholds (e.g., red for high response time, green for good) are critical. This allows users to quickly scan and understand the status of metrics without deep analysis, enabling rapid response to deviations.

  6. 6. A data analyst is designing a dashboard for a call center to monitor agent performance. The dashboard needs to display key metrics such as average call handling time, first call resolution rate, and customer satisfaction scores. To draw immediate attention to agents whose performance falls below predefined thresholds (e.g., call handling time > 5 minutes), which design element is MOST effective?

    Visualization

    • A. Including detailed data tables for each metric
    • B. Conditional formatting with color-coding
    • C. Using a minimalist design approach
    • D. Consistent chart types across all metrics
    Show answer

    B. Conditional formatting with color-coding

    Conditional formatting uses color or other visual cues to highlight data points that meet specific conditions or fall outside acceptable ranges. This allows users to quickly identify performance issues without manually scanning all data.

  7. 7. A data scientist is analyzing a customer feedback dataset. They notice that the 'Sentiment Score' column, which should range from -1 (negative) to 1 (positive), contains some values like -5, 2, and 'N/A'. To prepare this data for a machine learning model, which data cleansing technique would be most appropriate for handling the 'N/A' values?

    Data Mining

    • A. Deduplication
    • B. Data standardization
    • C. Outlier detection and removal
    • D. Imputation
    Show answer

    D. Imputation

    Imputation is the process of replacing missing data points (like 'N/A' in this case) with substituted values. This allows the dataset to remain complete for analysis and model training, rather than discarding rows with missing information.

  8. 8. A data engineer is designing an ETL pipeline for a new customer relationship management (CRM) system. The source system is an older, on-premise database, and the target is a cloud-based data warehouse. The business requires that all customer records, including new ones, updates to existing records, and deletions, are synchronized daily. Which data acquisition strategy is most efficient for capturing only the changes from the source system without transferring the entire dataset each day?

    Data Mining

    • A. Snapshot replication
    • B. Full database dump and reload
    • C. Manual data entry and validation
    • D. Change Data Capture (CDC)
    Show answer

    D. Change Data Capture (CDC)

    Change Data Capture (CDC) is specifically designed to identify and capture only the data that has changed in the source system since the last extraction. This is highly efficient for synchronizing data daily without the overhead of transferring the entire dataset, which is crucial for large databases.

  9. 9. A data scientist receives a dataset in a file format that uses a human-readable, plain-text structure where data items are separated by commas and each line represents a new record. The first line contains the column headers. What file format is being described?

    Data Concepts and Environments

    • A. YAML
    • B. CSV
    • C. XML
    • D. JSON
    Show answer

    B. CSV

    The description of data items separated by commas and each line being a new record, with the first line as headers, is the defining characteristic of a Comma Separated Values (CSV) file.

  10. 10. A data analyst is examining the distribution of customer ages. The data shows a few extremely old customers, causing the distribution to be skewed to the right. The analyst wants to understand the most typical age of a customer without being heavily influenced by these outliers. Which descriptive statistic should they use?

    Data Analysis

    • A. Range
    • B. Mean
    • C. Mode
    • D. Median
    Show answer

    D. Median

    When data is skewed or contains outliers, the mean can be heavily influenced and pulled in the direction of the skew/outliers, making it a less representative measure of the 'typical' value. The median, being the middle value, is resistant to extreme values and provides a better measure of central tendency in such cases.

  11. 11. A data analyst is querying a customer database to find all customers who have placed an order in the last 30 days but have not yet received a shipping confirmation. The database contains two tables: `Orders` (OrderID, CustomerID, OrderDate) and `Shipments` (ShipmentID, OrderID, ConfirmationDate). Which type of SQL join should the analyst use, combined with an appropriate WHERE clause, to achieve this result?

    Data Mining

    • A. FULL OUTER JOIN
    • B. INNER JOIN
    • C. LEFT JOIN
    • D. RIGHT JOIN
    Show answer

    C. LEFT JOIN

    A LEFT JOIN is required to return all records from the 'left' table (Orders) and the matching records from the 'right' table (Shipments). By then filtering for `Shipments.ConfirmationDate IS NULL`, you identify orders that have not yet received a shipping confirmation, while still retaining all orders from the Orders table.

  12. 12. A data analyst is designing a dashboard for a logistics company to track package delivery status. The dashboard needs to immediately alert dispatchers to any packages that are severely delayed, showing their current status (e.g., 'In Transit', 'Delayed', 'Delivered'). Instead of relying solely on text, the analyst wants to use visual cues to make the 'Delayed' status stand out prominently. Which design principle should the analyst apply to achieve this immediate visual alert?

    Visualization

    • A. Modularity
    • B. Data-Ink Ratio
    • C. Visual Cues (Color/Icons)
    • D. Gestalt Principles (Proximity)
    Show answer

    C. Visual Cues (Color/Icons)

    Using distinct colors (e.g., red) or warning icons for 'Delayed' status are visual cues that immediately draw attention and communicate urgency without requiring the user to read text, aligning with the goal of an immediate alert.

  13. 13. A data engineer is designing an ETL process to integrate data from a legacy system into a new data warehouse. The legacy system stores customer addresses in a single free-form text field, while the new data warehouse requires separate fields for street, city, state, and zip code. Which data transformation technique is most appropriate for handling this specific data structure change?

    Data Mining

    • A. Normalization
    • B. Filtering
    • C. Parsing
    • D. Aggregation
    Show answer

    C. Parsing

    Parsing is the most appropriate technique as it involves breaking down a complex data field into multiple, structured components. This directly addresses the need to separate the single address field into distinct street, city, state, and zip code fields.

  14. 14. A senior data analyst is reviewing a newly developed sales dashboard. The dashboard uses different color palettes for each region's sales charts (e.g., blue for North, green for South, red for East), and the titles for similar metrics vary across different tabs (e.g., 'Total Revenue' on one tab, 'Sales Sum' on another). While the individual charts are accurate, the analyst notes that users are struggling to quickly interpret information and switch between views. Which design principle is MOST likely being violated?

    Visualization

    • A. Consistency and Alignment
    • B. Proximity
    • C. Data-Ink Ratio
    • D. Visual Hierarchy
    Show answer

    A. Consistency and Alignment

    Inconsistent use of color palettes for the same type of data and varying terminology for identical metrics directly violates the principle of consistency, making the dashboard difficult to navigate and interpret.

  15. 15. A research institution is collecting genetic sequencing data, which consists of long strings of nucleotide bases (A, T, C, G) and associated metadata. This data is extremely large, does not fit neatly into rows and columns, and requires flexible storage that can accommodate varying structures and complex relationships without a rigid schema. Which database paradigm is most suitable for storing this type of data?

    Data Concepts and Environments

    • A. Relational Database
    • B. Key-Value Store
    • C. Graph Database
    • D. Document Database
    Show answer

    D. Document Database

    Genetic sequencing data, often stored as complex JSON-like objects or other flexible formats, benefits greatly from a document database. Document databases are designed to store semi-structured data, typically in JSON, BSON, or XML formats, allowing for flexible schemas and easy handling of nested data structures, which is ideal for genomic data with varying attributes and lengths.

  16. 16. During a customer data consolidation project, a master data management system identifies three source records for the same customer with conflicting phone numbers. The system is configured to automatically select the value from the most recently updated source system as the final value in the golden record. This configuration is an example of which MDM concept?

    Data Governance, Quality and Controls

    • A. Survivorship rule
    • B. Data lineage tracking
    • C. Referential integrity constraint
    • D. Data profiling rule
    Show answer

    A. Survivorship rule

    A survivorship rule determines which conflicting attribute value 'survives' into the golden record when multiple source systems provide different values for the same entity. Here, the rule is 'most recently updated wins.' Referential integrity concerns valid relationships between tables, lineage tracks data's history, and profiling analyzes data characteristics rather than resolving conflicts.

  17. 17. A data analyst is developing an internal dashboard for a software development team to track the progress of various features. The dashboard needs to clearly show the current status of each feature (e.g., 'To Do', 'In Progress', 'Testing', 'Done') and the lead developer assigned to it. The team lead needs to quickly identify bottlenecks and ensure workload distribution. Which visualization type is MOST effective for this scenario?

    Visualization

    • A. Bubble Chart
    • B. Scatter Plot
    • C. Gantt Chart
    • D. Pivot Table
    Show answer

    D. Pivot Table

    A pivot table (or matrix table) is highly effective for this scenario. It can display features as rows, status categories as columns, and individual cells can show the assigned developer. This allows for quick scanning of status per feature and, when aggregated, can reveal workload distribution and bottlenecks per developer or status category. While a Gantt chart shows timelines, it's less direct for current status and workload balance across multiple features and developers in a single, glanceable view that can be easily filtered or grouped.

  18. 18. A company assigns one employee to make business decisions about how customer data should be classified, who is authorized to access it, and how long it must be retained. A separate IT employee is responsible for implementing the backups, encryption, and storage infrastructure that protect that same data. Which role is accountable for the classification and access decisions?

    Data Governance, Quality and Controls

    • A. Data Custodian
    • B. Data Owner
    • C. Data Steward
    • D. System Administrator
    Show answer

    B. Data Owner

    The Data Owner is the accountable business role that makes decisions about classification, access authorization, and retention for a dataset. The Data Custodian executes the technical implementation (backups, encryption, storage) based on the owner's decisions, but does not set the classification or access policy itself.

  19. 19. A financial institution needs to store highly structured customer account information, including account numbers, balances, transaction history, and personal details. They require ACID properties to ensure data integrity during concurrent transactions. Which data environment is most appropriate for this requirement?

    Data Concepts and Environments

    • A. Relational Database
    • B. NoSQL Database
    • C. Data Warehouse
    • D. Data Lake
    Show answer

    A. Relational Database

    Relational databases are designed for highly structured data, enforce ACID properties for transactional integrity, and are ideal for managing consistent, reliable customer account information.

  20. 20. A business analyst is reviewing historical sales data to understand the long-term patterns and make future predictions. They observe that sales have been steadily increasing over the past five years, despite some seasonal fluctuations. To identify and quantify this underlying, consistent upward movement in sales, what aspect of trend analysis should they focus on?

    Data Analysis

    • A. Secular trend
    • B. Seasonal variation
    • C. Cyclical variation
    • D. Irregular variation
    Show answer

    A. Secular trend

    Secular trend, often simply referred to as 'trend', represents the long-term, underlying movement or direction in a time series. The observation of sales 'steadily increasing over the past five years' directly corresponds to identifying a secular trend.

  21. 21. A data analyst is examining a dataset of monthly sales figures for a retail company. They observe a general upward trend over several years, but also notice consistent peaks during holiday seasons and dips during off-peak months. Which component of time series data is primarily represented by the consistent peaks and dips occurring at regular intervals?

    Data Analysis

    • A. Cyclical
    • B. Trend
    • C. Irregular
    • D. Seasonality
    Show answer

    D. Seasonality

    Seasonality refers to predictable and recurrent patterns in time series data that occur over a fixed period, such as daily, weekly, monthly, or yearly. The consistent peaks during holiday seasons and dips during off-peak months perfectly describe seasonality.

  22. 22. A data analyst is preparing a visualization for a diverse audience, ranging from technical experts to non-technical business executives. The visualization needs to present complex statistical analysis results in a way that provides detailed information for the experts but also a clear, concise summary for the executives. Which design principle, if implemented effectively, would BEST address the needs of this mixed audience?

    Visualization

    • A. Data-Ink Ratio
    • B. Progressive Disclosure
    • C. Color Blindness Accessibility
    • D. Aesthetic Appeal
    Show answer

    B. Progressive Disclosure

    Progressive disclosure allows the visualization to initially present high-level summaries for a broad audience, with options for users to drill down into more detailed or technical information as needed, catering to both executives and experts.

  23. 23. A data analyst is working with a dataset containing customer feedback. Each entry includes a unique ID, the date of feedback submission, and a free-text comment describing the customer's experience. Which of the following best describes the data type of the 'free-text comment' field?

    Data Concepts and Environments

    • A. Boolean
    • B. Categorical
    • C. Numerical
    • D. Text
    Show answer

    D. Text

    The 'free-text comment' field contains unstructured human language, which is best classified as a text data type. This allows for storage of arbitrary strings of characters.

  24. 24. A data team is setting up a new data acquisition process for customer feedback forms. The forms are submitted through a web portal, generating JSON payloads. The team needs to ensure that, before any transformations or loading, the incoming JSON data consistently adheres to a predefined structure, including specific fields and data types. Which data acquisition step is primarily focused on verifying this structural consistency?

    Data Mining

    • A. Data profiling
    • B. Schema validation
    • C. Data storage
    • D. Data enrichment
    Show answer

    B. Schema validation

    Schema validation is the process of checking whether incoming data conforms to a predefined schema (structure, data types, constraints). For JSON payloads, this involves ensuring that the received data has all expected fields, and that their values are of the correct type, which is crucial early in the acquisition process.

  25. 25. A data governance committee is establishing policies for data acquisition from external vendors. They are particularly concerned about ensuring the data received is accurate, reliable, and adheres to privacy regulations. Which of the following aspects is MOST critical to address during the data acquisition phase to mitigate risks related to data quality and compliance?

    Data Mining

    • A. Establishing clear data contracts and SLAs with vendors.
    • B. Defining the schema for the target data warehouse.
    • C. Implementing advanced data visualization tools.
    • D. Optimizing the network bandwidth for data transfer.
    Show answer

    A. Establishing clear data contracts and SLAs with vendors.

    Establishing clear data contracts and Service Level Agreements (SLAs) with external vendors during the acquisition phase is paramount. These documents define data format, quality standards, refresh rates, security protocols, and compliance requirements (e.g., GDPR, CCPA), directly addressing accuracy, reliability, and privacy risks from the source.

CompTIA Data+ (DA0-002) flashcards

Tap a card to flip it. 149 flashcards in the full deck.

  • Grouped Bar Chart

    Flip card

    A bar chart that displays multiple sets of bars, grouped together for each category, allowing for direct comparison of sub-categories.

    • Compares multiple discrete categories.
    • Effective for showing performance across different groups.
    • Each group has its own set of bars representing sub-categories.
    Study this card →
  • Outlier Treatment (Winsorization/Normalization)

    Flip card

    Techniques used to manage the impact of extreme data points (outliers) on statistical analysis, either by transforming their values (winsorization) or by changing the data's scale (normalization).

    • Winsorization caps outliers at a specified percentile.
    • Normalization (e.g., log transform) can reduce skewness caused by outliers.
    • Aims to preserve data while reducing the distorting effect of extremes.
    Study this card →
  • Data Warehouse

    Flip card

    A large, centralized repository of integrated data from various disparate sources, optimized for analytical querying and reporting.

    • Stores historical and aggregated data, not real-time operational data.
    • Data is cleaned, transformed, and loaded (ETL) into a predefined schema (schema-on-write).
    • Used for business intelligence, trend analysis, and strategic decision-making.
    Study this card →
  • Dynamic Data Masking

    Flip card

    A security technique that obscures sensitive data in real time as it is queried, based on the user's role or privilege level, without modifying the underlying stored data.

    • Applied at query/view time, not to the stored data itself
    • Different users can see different masked/unmasked views of the same record
    • Contrasts with static masking, which permanently alters a data copy (e.g., for test/dev)
    Study this card →
  • Visual Cues (Thresholds)

    Flip card

    Graphical elements like color, size, or icons used to highlight data points that meet or exceed predefined thresholds, indicating immediate status or action.

    • Enables rapid scanning and comprehension.
    • Draws attention to critical data points.
    • Supports quick decision-making in dynamic environments.
    Study this card →
  • Conditional Formatting

    Flip card

    A visual design technique that applies specific formatting (e.g., color, icons, font styles) to data points in a visualization when certain conditions or rules are met.

    • Draws immediate attention to critical data points or outliers.
    • Helps users quickly identify trends, patterns, or anomalies.
    • Commonly used to highlight values above/below thresholds, or within specific ranges.
    Study this card →
  • Data Imputation

    Flip card

    Data imputation is the process of replacing missing data with substituted values. The goal is to fill in gaps in a dataset to maintain data integrity and enable complete analysis.

    • Common methods include mean, median, mode, or predictive imputation.
    • Helps maintain the dataset size and statistical power.
    • Choice of method depends on the nature of missing data and variable distribution.
    Study this card →
  • Change Data Capture (CDC)

    Flip card

    A set of software design patterns used to determine and track the data that has changed within a database since the last time data was extracted.

    • Enables incremental loading, significantly reducing data transfer volume and processing time.
    • Typically works by monitoring database transaction logs, using timestamps, or trigger-based mechanisms.
    • Essential for real-time or near real-time data synchronization in ETL/ELT pipelines.
    Study this card →
  • CSV (Comma Separated Values)

    Flip card

    A plain-text file format that stores tabular data in rows and columns, where each column value is separated by a comma.

    • Human-readable and simple to parse.
    • Often used for exchanging data between different applications.
    • Each line typically represents a data record, and values within a record are separated by a delimiter (usually a comma).
    Study this card →
  • Median's Robustness to Outliers

    Flip card

    The median is a measure of central tendency that is resistant to the influence of outliers and extreme values, making it a preferred statistic for skewed distributions.

    • Represents the middle value in an ordered dataset.
    • Not affected by the magnitude of extreme values, only their count.
    • Provides a more accurate 'typical' value for skewed data compared to the mean.
    Study this card →
  • SQL LEFT JOIN

    Flip card

    A type of SQL join that returns all records from the left table (first table in the FROM clause) and the matching records from the right table. If there is no match, the right side will contain NULL values.

    • Used to find records in one table that do or do not have corresponding records in another.
    • Preserves all rows from the 'left' table.
    • Often combined with `WHERE right_table.column IS NULL` to find unmatched records.
    Study this card →
  • Visual Cues (Color/Icons)

    Flip card

    The use of graphical elements like distinct colors, shapes, or icons to convey information, highlight important data, or draw attention to specific states or thresholds.

    • Provides immediate, non-textual communication.
    • Effectively highlights critical information or exceptions.
    • Should be used consistently and sparingly to avoid visual clutter.
    Study this card →
  • Data Parsing

    Flip card

    The process of breaking down a complex string or text field into multiple, structured components based on defined patterns, delimiters, or rules.

    • Used to extract meaningful information from unstructured or semi-structured data.
    • Commonly applied to address fields, log files, or free-form text.
    • Often involves regular expressions or specific parsing functions.
    Study this card →
  • Consistency and Alignment

    Flip card

    A design principle ensuring that similar visual elements (colors, fonts, labels, spacing) are treated in the same way across a visualization or dashboard, and that elements are neatly arranged.

    • Reduces cognitive load for users.
    • Improves readability and ease of understanding.
    • Applies to colors, fonts, terminology, layout, and spacing.
    Study this card →
  • Document Database

    Flip card

    A type of NoSQL database that stores data in flexible, semi-structured document formats (typically JSON, BSON, or XML), allowing for dynamic schemas and easy handling of hierarchical data.

    • Stores data as 'documents' (e.g., JSON objects).
    • Flexible schema (schema-on-read).
    • Ideal for semi-structured and unstructured data.
    Study this card →
  • Survivorship Rule (MDM)

    Flip card

    A predefined rule in master data management that determines which value 'survives' into the golden record when multiple source systems have conflicting data for the same attribute.

    • Common rules: most recent update, most trusted source, most complete value
    • Applied during the merge/match process when building golden records
    • Different from lineage, which just tracks data's origin/history
    Study this card →
  • Pivot Table (Matrix Table)

    Flip card

    A data summarization tool used to rearrange and aggregate data by one or more keys or dimensions, allowing for interactive exploration and analysis of large datasets.

    • Summarizes data from a larger table.
    • Allows for dynamic rearrangement (pivoting) of rows and columns.
    • Excellent for cross-tabulation and identifying patterns in categorical data.
    Study this card →
  • Data Owner vs. Data Custodian

    Flip card

    The Data Owner is the accountable business role deciding classification, access, and retention; the Data Custodian implements the technical controls (storage, backup, encryption) that enforce those decisions.

    • Owner = business accountability and decision-making authority
    • Custodian = technical implementation of security/storage controls
    • Steward = day-to-day data quality rule definition and enforcement
    Study this card →
  • Relational Database

    Flip card

    A type of database that stores and provides access to data points that are related to one another. Data is organized into tables (relations) with predefined schemas.

    • Uses SQL for data manipulation and querying.
    • Enforces ACID properties (Atomicity, Consistency, Isolation, Durability).
    • Ideal for structured data and transactional applications.
    Study this card →
  • Secular Trend

    Flip card

    The long-term, underlying movement or general direction (upward, downward, or stable) of a time series over an extended period, often spanning several years.

    • Represents the dominant long-term pattern.
    • Not influenced by short-term fluctuations like seasonality or random noise.
    • Can be linear or non-linear.
    Study this card →
  • Seasonality in Time Series

    Flip card

    A predictable and recurrent pattern in time series data that repeats over a fixed period, such as within a year, month, or week, often influenced by calendar events.

    • Repeats at fixed intervals (e.g., yearly, monthly).
    • Predictable and consistent.
    • Often driven by calendar events like holidays or weather.
    Study this card →
  • Progressive Disclosure

    Flip card

    A design principle where essential information is presented first, and more detailed or advanced information is revealed only when explicitly requested by the user.

    • Reduces information overload for new or general users.
    • Allows experienced users to access depth when needed.
    • Commonly implemented with drill-down features, expand/collapse sections, or tooltips.
    Study this card →
  • Text Data Type

    Flip card

    A data type used to store sequences of characters, often representing human language or descriptive information.

    • Can include letters, numbers, symbols, and spaces.
    • Often used for names, addresses, descriptions, and comments.
    • Typically requires more storage than numerical or boolean types.
    Study this card →
  • Schema Validation

    Flip card

    The process of verifying that a dataset's structure, data types, and constraints conform to a predefined schema or data model.

    • Crucial for maintaining data integrity at the point of ingestion.
    • Prevents malformed data from entering downstream systems.
    • Often uses tools like JSON Schema or XML Schema Definition (XSD).
    Study this card →

Questions are original practice items written to match the published exam objectives. Step2Study is not affiliated with or endorsed by any certification body.