CompTIA Data+ (DA0-002)Data AnalysisEasy
A data analyst is preparing a dataset for building a predictive model. They notice that one of the features, 'Income', has a wide range of values, from $20,000 to $500,000, while another feature, 'Number of Purchases', ranges from 1 to 50. Some machine learning algorithms are sensitive to features with different scales, potentially giving undue weight to features with larger numerical ranges. Which preprocessing technique should the analyst apply to address this issue?
- AData imputation
- BFeature scaling
- COne-hot encoding
- DPrincipal Component Analysis (PCA)
Show answer & explanationAnswer & explanation
Correct answer: B. Feature scaling
Feature scaling is a technique used to standardize or normalize the range of independent variables or features of data. This ensures that features with larger numerical ranges do not dominate the learning process of algorithms sensitive to scale, such as gradient descent-based algorithms or distance-based algorithms.
Why the other options are wrong
- A. Data imputation is used to fill in missing values in a dataset, which is a different problem from feature scaling.
- C. One-hot encoding is used for converting categorical variables into a numerical format, not for addressing scale differences in numerical features.
- D. Principal Component Analysis (PCA) is a dimensionality reduction technique, not primarily for standardizing feature ranges.
Feature Scaling
A data preprocessing technique used to standardize or normalize the range of independent variables (features) within a dataset. This prevents features with larger values from dominating the model's learning process.
- Standardizes or normalizes feature ranges.
- Important for algorithms sensitive to feature magnitudes (e.g., SVM, K-Means, Neural Networks).
- Methods include Min-Max scaling (normalization) and Standardization (Z-score scaling).
Memory trick: Clean, Scale, Transform, Encode: Data's ready to explode!