Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy
A data scientist is tasked with preparing a dataset for a machine learning model that will predict whether a customer will purchase a product. The dataset contains a feature called 'Customer_Segment' with nominal categorical values such as 'New Customer', 'Loyal Customer', and 'High-Value Customer'. Which data preprocessing technique should the data scientist use to convert this feature into a suitable format for most machine learning algorithms?
- ABinning
- BPrincipal Component Analysis (PCA)
- CFeature Scaling
- DOne-Hot Encoding
Show answer & explanationAnswer & explanation
Correct answer: D. One-Hot Encoding
One-Hot Encoding is the appropriate technique for converting nominal categorical data into a numerical format that machine learning algorithms can process without implying any ordinal relationship. Each category is transformed into a new binary feature.
Why the other options are wrong
- A. Binning converts numerical features into categorical bins, which is the opposite of the requirement.
- B. PCA is a dimensionality reduction technique for numerical data, not for encoding nominal categorical features.
- C. Feature Scaling adjusts the range of numerical features, not categorical ones.
One-Hot Encoding
A technique used to convert categorical variables into a numerical format that machine learning algorithms can understand. It creates new binary features for each category.
- Transforms nominal categorical data.
- Creates a new binary column for each unique category.
- Avoids implying ordinal relationships between categories.
Memory trick: Categorical data needs a 'hot' new look for the model.