AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A data scientist is building an AI/ML model to predict customer churn. The dataset includes a 'Subscription Tier' feature with values like 'Basic', 'Premium', and 'Enterprise'. To prepare this categorical feature for a machine learning algorithm, which technique is most appropriate if the model assumes numerical input and there is no inherent order among the tiers?
- ALabel Encoding
- BBinning
- COne-Hot Encoding
- DFeature Scaling
Show answer & explanationAnswer & explanation
Correct answer: C. One-Hot Encoding
One-Hot Encoding is the most appropriate technique for nominal categorical features (where there is no inherent order). It converts each category into a new binary feature, preventing the model from incorrectly inferring an ordinal relationship where none exists, which would happen with Label Encoding.
Why the other options are wrong
- A. Label Encoding assigns a unique integer to each category. This implies an ordinal relationship (e.g., 0 < 1 < 2) which is incorrect for 'Basic', 'Premium', 'Enterprise' without an inherent order, potentially misleading the model.
- B. Binning (or discretization) converts continuous numerical features into discrete categories (bins), which is not applicable here for an already categorical feature.
- D. Feature Scaling (e.g., normalization, standardization) is applied to numerical features to bring them to a similar range, not to convert categorical features.
One-Hot Encoding
A technique to convert categorical variables into a numerical format that machine learning algorithms can understand, by creating binary columns for each category.
- Used for nominal (unordered) categorical data.
- Creates N new binary features for N categories.
- Avoids implying an artificial ordinal relationship.
- Can lead to increased dimensionality ('curse of dimensionality').
Memory trick: Categories need numbers, but order matters.