Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureMedium
A data scientist is preparing a dataset that contains a 'Product_Category' feature with values like 'Electronics', 'Clothing', 'Home Goods', and 'Books'. These categories have no inherent order or numerical relationship. To use this feature in a machine learning model, it must be converted into a numerical format without implying any false sense of order or magnitude. Which technique is most appropriate?
- AOne-Hot Encoding
- BFeature Scaling
- CTarget Encoding
- DLabel Encoding
Show answer & explanationAnswer & explanation
Correct answer: A. One-Hot Encoding
One-Hot Encoding is the most appropriate technique for nominal categorical features like 'Product_Category'. It creates new binary features for each category, preventing the model from assuming an ordinal relationship or magnitude that doesn't exist.
Why the other options are wrong
- B. Feature Scaling is for numerical features to adjust their range or distribution, not for converting categorical features.
- C. Target Encoding replaces a categorical value with the mean of the target variable for that category, which can be useful but is different from representing the categories themselves for a model.
- D. Label Encoding assigns a unique integer to each category (e.g., 'Electronics'=1, 'Clothing'=2), which implies an ordinal relationship that doesn't exist, potentially misleading the model.
One-Hot Encoding
A technique used to convert categorical variables into a numerical format that machine learning algorithms can understand, by creating new binary features for each category.
- Ideal for nominal (unordered) categorical features.
- Avoids implying ordinal relationships or magnitudes.
- Can lead to a high-dimensional dataset if many categories exist.
Memory trick: Categories to Numbers: Pick Your Converter