Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureHard

A data engineering team is tasked with preparing a large dataset for a machine learning model. The dataset contains several features with widely different ranges, such as 'Age' (0-100) and 'Income' (10,000-1,000,000). Many machine learning algorithms, especially those based on distance calculations like K-Nearest Neighbors or Support Vector Machines, perform poorly with such disparities. Which preprocessing technique should be applied to address this issue?

  1. AFeature scaling
  2. BDimensionality reduction
  3. COne-hot encoding
  4. DData imputation
Show answer & explanation

Correct answer: A. Feature scaling

Feature scaling (e.g., normalization or standardization) is essential when features have different ranges. It transforms features to a common scale, preventing features with larger values from dominating algorithms that rely on distance calculations, thus improving model performance.

Why the other options are wrong

  • B. Dimensionality reduction reduces the number of features, not specifically addressing the issue of differing numerical ranges of existing features.
  • C. One-hot encoding converts categorical data to numerical, not for numerical range disparities.
  • D. Data imputation fills missing values, unrelated to feature range differences.

Feature Scaling

A data preprocessing technique used to standardize or normalize the range of independent variables (features) in a dataset.

  • Prevents features with larger magnitudes from dominating distance-based algorithms.
  • Common methods: Min-Max Scaling (Normalization) and Standardization (Z-score normalization).
  • Essential for algorithms like K-Nearest Neighbors, Support Vector Machines, and neural networks.

Memory trick: Feature Scaling 'levels' the 'playing field' for numerical data.

More Describe fundamental principles of machine learning on Azure questions