Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy
A machine learning engineer needs to prepare a dataset where one of the features, 'City', contains nominal categorical data (e.g., 'New York', 'London', 'Tokyo'). This feature needs to be converted into a numerical format that a machine learning model can process without implying any ordinal relationship between the cities. Which data transformation technique should be applied?
- AMin-Max Scaling
- BStandardization
- COne-Hot Encoding
- DLabel Encoding
Show answer & explanationAnswer & explanation
Correct answer: C. One-Hot Encoding
One-Hot Encoding is the appropriate technique for converting nominal categorical data into a numerical format without implying any order or numerical relationship. It creates new binary features for each category.
Why the other options are wrong
- A. Min-Max Scaling is a feature scaling technique for numerical data, not categorical conversion.
- B. Standardization (Z-score normalization) is a feature scaling technique for numerical data, not categorical conversion.
- D. Label Encoding assigns a unique integer to each category, which can imply an ordinal relationship if not handled carefully.
One-Hot Encoding
One-Hot Encoding is a process of converting categorical variables into a numerical format that machine learning algorithms can understand. It creates new binary columns for each category, where '1' indicates the presence of that category and '0' indicates its absence.
- Used for nominal categorical data (no inherent order).
- Prevents the model from assuming ordinal relationships.
- Can lead to a high-dimensional sparse dataset if many categories exist.
Memory trick: One-Hot: Each category gets its 'one hot' spot in a new column, like a unique switch.