CompTIA Data+ (DA0-002) practice questions

230 free questions with answers and explanations.

Practice test
  1. 101.A data analyst is working with a large dataset of customer demographics and purchase history. They want to identify distinct groups of customers based on their similarities in age, income, and purchasing behavior without any prior knowledge of what these groups might be. Which unsupervised machine learning technique is best suited for this task?Data Analysis
  2. 102.A data analyst is comparing the performance of two different website layouts (Layout A and Layout B) by conducting an A/B test. They randomly assign users to one of the layouts and measure the average time spent on the page. After collecting data, they want to determine if there is a statistically significant difference in average time spent between the two layouts. Which inferential statistical test should they use?Data Analysis
  3. 103.A data scientist is working with a dataset containing customer transaction IDs. They notice that the ID column, which should be unique for each transaction, contains several duplicate entries. If these duplicates are not removed, they could lead to incorrect aggregations and skewed analysis. What is the most direct and effective data cleansing technique to address this issue?Data Mining
  4. 104.A data analyst is working with a dataset of product ratings from customers, where ratings are on a scale of 1 to 10. They want to understand the typical rating, but they notice that the dataset has several extreme outliers—a few very low ratings and a few very high ratings that seem to be anomalies. If the analyst wants a measure of central tendency that is least affected by these outliers, which measure should they choose?Data Analysis
  5. 105.A data engineering team is designing a new data storage solution that needs to handle high volumes of rapidly changing, schema-less data, primarily for real-time analytics and caching. Which of the following database types would be MOST suitable for this requirement?Data Concepts and Environments
  6. 106.A business intelligence analyst is building a report that requires combining customer demographic data with their sales transaction history. They have two tables: `Customers` (CustomerID, Name, Address) and `Transactions` (TransactionID, CustomerID, Product, Amount). They only want to see customers who have made at least one transaction and the details of those specific transactions. Which SQL JOIN type should be used?Data Mining
  7. 107.A data scientist is analyzing a large dataset of customer transactions. They want to group customers based on their purchasing behavior without any prior knowledge of customer segments. The goal is to discover natural groupings within the data. Which unsupervised learning technique is most appropriate for this task?Data Analysis
  8. 108.A data team is presenting a quarterly business review to senior leadership. The presentation needs to highlight the most significant achievements and challenges, supported by data, and conclude with strategic recommendations. Which of the following storytelling techniques should be prioritized to ensure the presentation is impactful and drives decision-making?Visualization
  9. 109.A data analyst is presented with a dataset of customer transaction values. The values are: $25, $30, $35, $40, $45, $50, $55, $60, $65, $70. What is the median of this dataset?Data Analysis
  10. 110.To reduce the risk of fraud, a company requires that the employee who initiates a vendor payment record cannot be the same employee who approves that payment for processing. This control is an example of which governance principle?Data Governance, Quality and Controls
  11. 111.A data team is presenting the results of a complex A/B test to a diverse group of stakeholders, including product managers, engineers, and marketing specialists. The presentation needs to convey the statistical significance of the results, the impact on user behavior, and the recommended next steps. Which design principle should be applied to bridge the understanding gap between technical details and business implications?Visualization
  12. 112.A company is developing an AI model to predict customer churn. The model will analyze historical customer data, including demographics, service usage patterns, and past interactions. The primary goal is to identify customers at risk of leaving so targeted retention campaigns can be launched. Which type of AI task does this scenario best represent?Data Concepts and Environments
  13. 113.A data scientist is preparing a presentation to explain the impact of a new marketing campaign to non-technical business stakeholders. The presentation needs to clearly show the increase in website traffic and conversion rates post-campaign launch. Which storytelling element should the data scientist prioritize to make the data more relatable and memorable for this audience?Visualization
  14. 114.A data analyst is designing a dashboard to monitor the performance of a company's sales representatives. The dashboard needs to clearly show the sales volume generated by each representative, allowing managers to quickly identify top performers and those needing improvement. Additionally, the dashboard should display the percentage contribution of each representative to the total sales. Which chart type would be MOST effective for visualizing both individual performance and proportional contribution to the total?Visualization
  15. 115.A company wants to grant database access based on employees' job functions, such as 'Sales Analyst' or 'HR Manager', rather than assigning permissions individually to each user. New employees automatically receive the correct permissions when assigned to a function. Which access control model should the company implement?Data Governance, Quality and Controls
  16. 116.A data analyst is evaluating the effectiveness of a new customer service training program. They collect customer satisfaction scores (on a scale of 1 to 10) from two groups: one group whose representatives underwent the new training and another group whose representatives did not. The analyst wants to determine if there is a statistically significant difference in satisfaction scores between the two groups, assuming the data is not normally distributed and the sample sizes are small. Which statistical test should the analyst use?Data Analysis
  17. 117.A university research team publishes a dataset of survey responses for public use. Before release, the team strips all direct identifiers, aggregates responses into broad age and income brackets, and removes any combination of fields that could indirectly re-identify a participant, such that even the research team itself cannot trace a response back to an individual. Which technique is being applied?Data Governance, Quality and Controls
  18. 118.A market researcher is conducting a survey to estimate the average spending of customers in a particular demographic. They collect data from a sample and calculate a 95% confidence interval for the mean spending. This interval is ($120, $150). What does this interval imply?Data Analysis
  19. 119.A healthcare provider is implementing a new electronic health record (EHR) system. The system needs to classify patient data into categories such as 'confidential', 'sensitive', and 'public' based on regulatory requirements (e.g., HIPAA). This classification is crucial for controlling access and ensuring compliance. Which data concept is being applied here?Data Concepts and Environments
  20. 120.A data analyst is working with a dataset of customer ages. The ages are heavily skewed to the right due to a small number of very old customers. Which measure of central tendency would be most appropriate to describe the typical customer age in this scenario?Data Analysis
  21. 121.A data governance committee is defining standards for data acquisition from external vendors. They require that all incoming datasets must adhere to a predefined structure, including specific column names, data types, and nullability constraints, to ensure compatibility with their existing systems. This agreement, specifying the format and characteristics of the data to be exchanged, is best described as a:Data Mining
  22. 122.A data quality specialist is auditing a customer database and finds that the 'Email' column sometimes contains 'NULL' values, empty strings (''), or even incorrect formats like 'not_available' or 'customer@domain'. They need to identify all records where the email address is effectively missing or invalid for a new marketing campaign. Which SQL query construct would be most suitable to identify these records?Data Mining
  23. 123.A data analyst is performing an exploratory data analysis on a sales dataset. They have two tables: `Customers` (CustomerID, CustomerName, City) and `Orders` (OrderID, CustomerID, OrderDate, Amount). The analyst needs a list of ALL customers, including those who have placed no orders, AND all orders, even those without a matching customer (due to data entry errors). Which SQL JOIN type should be used?Data Mining
  24. 124.A data engineer is designing a data pipeline to process customer interaction data from a web application. The data is generated in real-time and includes user actions, timestamps, and session identifiers. This data needs to be stored in a way that allows for immediate, high-volume ingestion and fast retrieval of individual records based on their session ID. Which database type is most appropriate for this scenario?Data Concepts and Environments
  25. 125.A data scientist is preparing a dataset for a predictive model where the 'Age' column contains several entries like 'twenty-five', '30 years old', and '45'. Before converting these to a numerical format, which data cleansing step is most critical to ensure consistency and prevent errors during the conversion?Data Mining
  26. 126.A data analyst is evaluating customer feedback scores. The scores range from 1 to 10, and the analyst wants to understand the central tendency, but the data is heavily skewed with many low scores and a few very high scores. Which measure of central tendency would be LEAST appropriate for representing the typical score in this dataset?Data Analysis
  27. 127.A data scientist is preparing to conduct a hypothesis test to determine if a new website design leads to a statistically significant increase in conversion rates. Before collecting data, they need to establish the maximum acceptable probability of incorrectly rejecting a true null hypothesis. What statistical concept are they defining?Data Analysis
  28. 128.A data engineering team is setting up a new data pipeline for their e-commerce platform. They need to move raw transactional data from a production database into a data lake, then clean and transform it for analytical reporting in a data warehouse. The process involves ingesting data as-is, then performing complex transformations and aggregations later. Which data integration approach is best suited for this scenario?Data Mining
  29. 129.A data engineer is designing a system to store sensor data from IoT devices. Each sensor periodically transmits readings that include a timestamp, device ID, and temperature. The engineer anticipates a very high volume of data, with frequent writes and less frequent reads, primarily for time-series analysis. Data consistency is important, but absolute ACID compliance for every single reading is not a strict requirement, prioritizing write availability and scalability. Which database type is most suitable?Data Concepts and Environments
  30. 130.A file-sharing platform allows each document's creator to individually decide which other users may view, edit, or share that specific document, granting or revoking permissions at their own discretion without requiring administrator approval. Which access control model does this describe?Data Governance, Quality and Controls
  31. 131.A quality control manager is monitoring the weight of cereal boxes coming off an assembly line. The target weight is 450 grams. A sample of 30 boxes is collected, and their average weight is 448 grams with a standard deviation of 5 grams. The manager wants to determine if the average weight of cereal boxes is significantly different from the target weight. Assuming the population standard deviation is unknown, which statistical test should be used?Data Analysis
  32. 132.A data analyst is examining sales data for a new product launched three months ago. The daily sales figures show a clear upward trajectory, but with significant day-to-day fluctuations. The analyst wants to understand the underlying growth pattern, free from the daily noise. Which analytical technique would best help to identify this underlying growth?Data Analysis
  33. 133.A data engineer is designing an ETL process for a new enterprise data warehouse. The source system contains customer data from various legacy applications, each with slightly different data formats and schemas. The engineer needs to ensure that customer records from all sources are merged into a single, consistent format before being loaded into the data warehouse. Which phase of the ETL process is primarily responsible for unifying these disparate data structures and preparing them for the target system?Data Mining
  34. 134.A data analyst is working with a sales dataset that includes a 'TransactionAmount' column. They notice that a few entries are exceptionally high, skewing statistical measures like the mean and standard deviation. These extreme values are legitimate but rare. To mitigate their impact on statistical analysis without removing them entirely, the analyst decides to cap these values at the 99th percentile. What is this data cleansing technique called?Data Mining
  35. 135.A data analyst is designing a financial dashboard for investment managers. The dashboard must clearly show the performance of several investment portfolios over the past five years, with each portfolio's performance displayed as a continuous trend. Additionally, the managers need to quickly identify periods of significant volatility or divergence among the portfolios. Which chart type is MOST effective for this scenario?Visualization
  36. 136.A data analyst is working with a large dataset of customer demographics and purchase history. They want to identify distinct groups of customers based on their purchasing behavior (e.g., high-value frequent buyers, occasional bargain hunters, loyal brand advocates). The analyst does not have pre-defined labels for these customer segments but wants to discover natural groupings within the data. Which unsupervised machine learning technique is most suitable for this task?Data Analysis
  37. 137.A data quality team audits 2,500 customer records against a trusted reference dataset to verify that field values correctly represent real-world facts. The audit finds that 2,275 records match the reference data exactly. Which value represents the accuracy rate of this dataset?Data Governance, Quality and Controls
  38. 138.A data architect is evaluating different storage solutions for a new social media analytics platform. The platform needs to store user profiles, posts, comments, and the intricate connections between users (e.g., followers, friends, mentions). The primary requirement is to efficiently query and analyze these complex, multi-directional relationships. Which database paradigm would be most effective for this task?Data Concepts and Environments
  39. 139.A data engineer is designing an ETL process for a new data warehouse. They have identified several source systems, each with its own schema and data types. Before loading the data into the target warehouse, the engineer needs to ensure that all incoming data conforms to a predefined set of rules and data formats to maintain data quality and consistency. Which of the following phases of the ETL process is primarily responsible for this activity?Data Mining
  40. 140.A data analyst is preparing a dataset of customer feedback where survey responses include a free-text field for 'Suggestions'. Many entries contain leading or trailing spaces, multiple spaces between words, or unwanted newline characters. To clean this text for consistent analysis, which combination of SQL string functions would be most effective?Data Mining
  41. 141.A data analyst calculates the mean score of a certification exam to be 75 with a standard deviation of 10. Assuming the scores are normally distributed, what percentage of test-takers would be expected to score between 65 and 85?Data Analysis
  42. 142.A data engineer is tasked with building a robust data pipeline that sources customer data from multiple regional databases. Each database uses a slightly different identifier for customers (e.g., `customer_id`, `cust_num`, `client_id`). To create a unified customer view, the engineer needs to convert these various identifiers into a single, consistent `global_customer_id` format. This process also involves mapping legacy codes to new standardized codes. Which data transformation technique is being primarily applied here?Data Mining
  43. 143.A team of data analysts is investigating the relationship between advertising spend and sales revenue. They have collected data over several months and want to quantify the strength and direction of the linear relationship between these two continuous variables. Which statistical measure should they calculate?Data Analysis
  44. 144.A data analyst is preparing a quarterly business review for senior leadership, focusing on the performance of various marketing campaigns. The presentation needs to convey the impact of each campaign on sales, clearly showing which campaigns generated the most revenue and which were less effective. The analyst also wants to display the budget allocated to each campaign alongside its revenue. Which chart type would be MOST effective for comparing multiple campaigns' revenue and budget simultaneously?Visualization
  45. 145.A financial analyst is comparing the volatility of two different stock portfolios. Portfolio A has a standard deviation of returns of 12%, while Portfolio B has a standard deviation of returns of 8%. Both portfolios have the same average return. What can the analyst infer from these standard deviations?Data Analysis
  46. 146.A data analyst is examining a dataset of daily website visitors. They notice that the number of visitors tends to be lower on weekends and higher on weekdays, and there's a general upward trend over the year. To isolate and understand these regular, predictable movements within a specific time period, what component of time series analysis should the analyst focus on?Data Analysis
  47. 147.A data analyst is examining a dataset of employee salaries. The salaries are known to be right-skewed, meaning a few highly paid executives significantly inflate the average. The analyst wants to report a measure of central tendency that is resistant to these extreme values, providing a better representation of what a 'typical' employee earns. Which of the following descriptive statistics should they use?Data Analysis
  48. 148.A data analyst is performing an exploratory data analysis on a customer feedback dataset. They observe that the 'Feedback_Text' column is often filled with short, informal phrases and sometimes includes emoticons and repetitive punctuation (e.g., 'Great!!!', 'So happy :)'). Before applying sentiment analysis, these elements need to be handled. Which data cleansing technique specifically targets removing or standardizing such informal text features?Data Mining
  49. 149.A company is developing an AI-powered customer service chatbot. To ensure the chatbot provides unbiased and equitable responses across different user demographics, the development team needs to continuously monitor its performance for potential biases in sentiment analysis and response generation. Which AI concept are they primarily focusing on?Data Concepts and Environments
  50. 150.A data quality specialist is analyzing a dataset of customer feedback where users are asked to rate their experience on a scale of 1 to 5. They discover that some entries contain values like '6', 'Excellent', or 'N/A'. To prepare this data for numerical analysis, these invalid entries must be addressed. Which data cleansing technique is most appropriate to ensure all values conform to the expected 1-5 numerical scale?Data Mining