AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy

A data scientist is building an AI/ML model to predict customer churn. The dataset includes a 'Subscription Tier' feature with values like 'Basic', 'Premium', and 'Enterprise'. To prepare this categorical feature for a machine learning algorithm, which technique is most appropriate if the model assumes numerical input and there is no inherent order among the tiers?

  1. ALabel Encoding
  2. BBinning
  3. COne-Hot Encoding
  4. DFeature Scaling
Show answer & explanation

Correct answer: C. One-Hot Encoding

One-Hot Encoding is the most appropriate technique for nominal categorical features (where there is no inherent order). It converts each category into a new binary feature, preventing the model from incorrectly inferring an ordinal relationship where none exists, which would happen with Label Encoding.

Why the other options are wrong

  • A. Label Encoding assigns a unique integer to each category. This implies an ordinal relationship (e.g., 0 < 1 < 2) which is incorrect for 'Basic', 'Premium', 'Enterprise' without an inherent order, potentially misleading the model.
  • B. Binning (or discretization) converts continuous numerical features into discrete categories (bins), which is not applicable here for an already categorical feature.
  • D. Feature Scaling (e.g., normalization, standardization) is applied to numerical features to bring them to a similar range, not to convert categorical features.

One-Hot Encoding

A technique to convert categorical variables into a numerical format that machine learning algorithms can understand, by creating binary columns for each category.

  • Used for nominal (unordered) categorical data.
  • Creates N new binary features for N categories.
  • Avoids implying an artificial ordinal relationship.
  • Can lead to increased dimensionality ('curse of dimensionality').

Memory trick: Categories need numbers, but order matters.

More AI/ML and Generative AI Fundamentals questions