Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy

A data scientist is preparing a dataset for a machine learning model that predicts house prices. The dataset contains a 'Neighborhood' column with categorical values like 'Downtown', 'Suburban', and 'Rural'. To make this data suitable for most machine learning algorithms, which preprocessing technique should be applied?

  1. AFeature scaling
  2. BNormalization
  3. CData imputation
  4. DOne-hot encoding
Show answer & explanation

Correct answer: D. One-hot encoding

One-hot encoding converts categorical variables into a numerical format that machine learning algorithms can process without implying any ordinal relationship between categories. This is crucial for nominal categories like 'Neighborhood'.

Why the other options are wrong

  • A. Feature scaling adjusts the range of numerical features, similar to normalization, and is not for categorical data.
  • B. Normalization scales numerical features to a standard range, which is not applicable to categorical data directly.
  • C. Data imputation is used to fill in missing values, not to convert categorical data into a numerical format.

One-hot Encoding

A technique that converts categorical variables into a numerical format, where each category is represented by a binary vector (0s and 1s).

  • Used for nominal categorical data (no inherent order).
  • Creates new binary columns for each unique category.
  • Prevents algorithms from assuming false ordinal relationships.

Memory trick: Categorical data needs a 'hot' new look for ML.

More Describe fundamental principles of machine learning on Azure questions