Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureMedium

A data scientist is preparing a dataset that contains a 'Product_Category' feature with values like 'Electronics', 'Clothing', 'Home Goods', and 'Books'. These categories have no inherent order or numerical relationship. To use this feature in a machine learning model, it must be converted into a numerical format without implying any false sense of order or magnitude. Which technique is most appropriate?

  1. AOne-Hot Encoding
  2. BFeature Scaling
  3. CTarget Encoding
  4. DLabel Encoding
Show answer & explanation

Correct answer: A. One-Hot Encoding

One-Hot Encoding is the most appropriate technique for nominal categorical features like 'Product_Category'. It creates new binary features for each category, preventing the model from assuming an ordinal relationship or magnitude that doesn't exist.

Why the other options are wrong

  • B. Feature Scaling is for numerical features to adjust their range or distribution, not for converting categorical features.
  • C. Target Encoding replaces a categorical value with the mean of the target variable for that category, which can be useful but is different from representing the categories themselves for a model.
  • D. Label Encoding assigns a unique integer to each category (e.g., 'Electronics'=1, 'Clothing'=2), which implies an ordinal relationship that doesn't exist, potentially misleading the model.

One-Hot Encoding

A technique used to convert categorical variables into a numerical format that machine learning algorithms can understand, by creating new binary features for each category.

  • Ideal for nominal (unordered) categorical features.
  • Avoids implying ordinal relationships or magnitudes.
  • Can lead to a high-dimensional dataset if many categories exist.

Memory trick: Categories to Numbers: Pick Your Converter

More Describe fundamental principles of machine learning on Azure questions