CompTIA Data+ (DA0-002)Data MiningEasy
A data analyst is preparing a dataset for a new machine learning model. They notice that several records have missing values in a critical 'Customer_Age' column, which is essential for the model's accuracy. Which of the following data cleansing techniques is most appropriate for handling these missing values if the analyst wants to preserve as much data as possible while providing a reasonable estimate?
- AReplacing missing 'Customer_Age' values with a constant, such as '0' or 'Unknown'.
- BDeleting all records that contain any missing values.
- CIgnoring the missing values and allowing the machine learning model to handle them directly.
- DImputing the missing 'Customer_Age' values with the mean age of all existing customers.
Show answer & explanationAnswer & explanation
Correct answer: D. Imputing the missing 'Customer_Age' values with the mean age of all existing customers.
Imputing missing values with the mean (or median for skewed data) is a common and effective technique to fill gaps in numerical data while retaining the majority of the dataset. This approach provides a statistically sound estimate without losing valuable records.
Why the other options are wrong
- A. Replacing with a constant like '0' or 'Unknown' can introduce bias and artificial patterns into the data, misleading the machine learning model, especially for numerical columns.
- B. Deleting records with missing values can lead to significant data loss, especially if missingness is widespread, potentially biasing the dataset and reducing model performance.
- C. Ignoring missing values is often not an option, as many machine learning models cannot process NaN (Not a Number) values directly and will either error out or produce unreliable results.
Mean Imputation
A data cleansing technique where missing numerical values in a dataset are replaced with the average (mean) of the available values for that specific feature.
- Preserves data volume by not deleting rows.
- Introduces less bias than deleting rows if data is missing at random.
- Can reduce variance and distort relationships if used excessively or on non-random missing data.
Memory trick: Missing pieces need a smart fill, not a complete spill.