Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy
A data scientist is preparing a dataset for a machine learning model that predicts house prices. The dataset contains a 'Neighborhood' column with categorical values like 'Downtown', 'Suburban', and 'Rural'. To make this data suitable for most machine learning algorithms, which preprocessing technique should be applied?
- AFeature scaling
- BNormalization
- CData imputation
- DOne-hot encoding
Show answer & explanationAnswer & explanation
Correct answer: D. One-hot encoding
One-hot encoding converts categorical variables into a numerical format that machine learning algorithms can process without implying any ordinal relationship between categories. This is crucial for nominal categories like 'Neighborhood'.
Why the other options are wrong
- A. Feature scaling adjusts the range of numerical features, similar to normalization, and is not for categorical data.
- B. Normalization scales numerical features to a standard range, which is not applicable to categorical data directly.
- C. Data imputation is used to fill in missing values, not to convert categorical data into a numerical format.
One-hot Encoding
A technique that converts categorical variables into a numerical format, where each category is represented by a binary vector (0s and 1s).
- Used for nominal categorical data (no inherent order).
- Creates new binary columns for each unique category.
- Prevents algorithms from assuming false ordinal relationships.
Memory trick: Categorical data needs a 'hot' new look for ML.