Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy
A data scientist is preparing a dataset for a machine learning model. The dataset contains a feature named 'ProductCategory' with values like 'Electronics', 'Clothing', 'Home Goods', and 'Books'. To ensure the model can correctly interpret these categorical values without implying any ordinal relationship, which data preprocessing technique should be applied?
- AFeature Scaling
- BOne-Hot Encoding
- CPrincipal Component Analysis (PCA)
- DBinning
Show answer & explanationAnswer & explanation
Correct answer: B. One-Hot Encoding
One-Hot Encoding is the appropriate technique for converting categorical variables into a numerical format that machine learning models can process, without introducing an artificial order. It creates new binary columns for each category.
Why the other options are wrong
- A. Feature Scaling adjusts the range of features, which is not the primary goal for categorical data without inherent order.
- C. Principal Component Analysis (PCA) is a dimensionality reduction technique for numerical data, not for encoding categorical features.
- D. Binning groups continuous numerical data into bins, which is not applicable here as 'ProductCategory' is already categorical.
One-Hot Encoding
One-Hot Encoding is a process of converting categorical variable values into a numerical format that machine learning algorithms can understand. It creates new binary features for each unique category.
- Transforms nominal categorical data.
- Creates a new binary column for each category.
- Prevents the model from assuming an ordinal relationship.
- Can increase the dimensionality of the dataset.
Memory trick: Categories need a 'hot' new look to be understood by the machine.