Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy

A machine learning engineer needs to prepare a dataset where one of the features, 'City', contains nominal categorical data (e.g., 'New York', 'London', 'Tokyo'). This feature needs to be converted into a numerical format that a machine learning model can process without implying any ordinal relationship between the cities. Which data transformation technique should be applied?

  1. AMin-Max Scaling
  2. BStandardization
  3. COne-Hot Encoding
  4. DLabel Encoding
Show answer & explanation

Correct answer: C. One-Hot Encoding

One-Hot Encoding is the appropriate technique for converting nominal categorical data into a numerical format without implying any order or numerical relationship. It creates new binary features for each category.

Why the other options are wrong

  • A. Min-Max Scaling is a feature scaling technique for numerical data, not categorical conversion.
  • B. Standardization (Z-score normalization) is a feature scaling technique for numerical data, not categorical conversion.
  • D. Label Encoding assigns a unique integer to each category, which can imply an ordinal relationship if not handled carefully.

One-Hot Encoding

One-Hot Encoding is a process of converting categorical variables into a numerical format that machine learning algorithms can understand. It creates new binary columns for each category, where '1' indicates the presence of that category and '0' indicates its absence.

  • Used for nominal categorical data (no inherent order).
  • Prevents the model from assuming ordinal relationships.
  • Can lead to a high-dimensional sparse dataset if many categories exist.

Memory trick: One-Hot: Each category gets its 'one hot' spot in a new column, like a unique switch.

More Describe fundamental principles of machine learning on Azure questions