CompTIA Data+ (DA0-002)Data AnalysisMedium
A data scientist is building a predictive model for housing prices. They have a dataset with various features such as square footage, number of bedrooms, and location. Before feeding this data into a machine learning algorithm, they want to ensure that all numerical features contribute equally to the model, regardless of their original scale. For example, square footage might range from 500 to 5000, while the number of bedrooms might range from 1 to 6. What data preprocessing technique should they apply?
- AFeature scaling
- BDimensionality reduction
- COne-hot encoding
- DImputation
Show answer & explanationAnswer & explanation
Correct answer: A. Feature scaling
Feature scaling (e.g., normalization or standardization) is a crucial preprocessing step that transforms numerical features to a common scale without distorting differences in the ranges of values or losing information. This ensures that features with larger numerical ranges do not disproportionately influence the model's learning process.
Why the other options are wrong
- B. Dimensionality reduction techniques reduce the number of features, not necessarily their scale.
- C. One-hot encoding is used for converting categorical variables into a numerical format.
- D. Imputation is used to fill in missing values in a dataset.
Feature Scaling
A data preprocessing technique used to transform numerical features to a standard range or distribution, preventing features with larger values from dominating the learning process of machine learning algorithms.
- Includes methods like normalization (Min-Max Scaling) and standardization (Z-score Scaling).
- Essential for algorithms sensitive to feature scales (e.g., K-Nearest Neighbors, Support Vector Machines).
- Helps improve model convergence and performance.
Memory trick: Feature Scaling is like putting all your ingredients on a kitchen scale so they're measured fairly.