Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureHard
A data engineering team is tasked with preparing a large dataset for a machine learning model. The dataset contains several features with widely different ranges, such as 'Age' (0-100) and 'Income' (10,000-1,000,000). Many machine learning algorithms, especially those based on distance calculations like K-Nearest Neighbors or Support Vector Machines, perform poorly with such disparities. Which preprocessing technique should be applied to address this issue?
- AFeature scaling
- BDimensionality reduction
- COne-hot encoding
- DData imputation
Show answer & explanationAnswer & explanation
Correct answer: A. Feature scaling
Feature scaling (e.g., normalization or standardization) is essential when features have different ranges. It transforms features to a common scale, preventing features with larger values from dominating algorithms that rely on distance calculations, thus improving model performance.
Why the other options are wrong
- B. Dimensionality reduction reduces the number of features, not specifically addressing the issue of differing numerical ranges of existing features.
- C. One-hot encoding converts categorical data to numerical, not for numerical range disparities.
- D. Data imputation fills missing values, unrelated to feature range differences.
Feature Scaling
A data preprocessing technique used to standardize or normalize the range of independent variables (features) in a dataset.
- Prevents features with larger magnitudes from dominating distance-based algorithms.
- Common methods: Min-Max Scaling (Normalization) and Standardization (Z-score normalization).
- Essential for algorithms like K-Nearest Neighbors, Support Vector Machines, and neural networks.
Memory trick: Feature Scaling 'levels' the 'playing field' for numerical data.