AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy

A data scientist is preparing a dataset for an AI/ML model that will predict customer churn. One of the features is 'Subscription Type', which can be 'Basic', 'Premium', or 'Enterprise'. Which of the following data preprocessing techniques is most appropriate for this categorical feature if the model expects numerical input and there is no inherent order among the subscription types?

  1. AOne-Hot Encoding
  2. BMin-Max Scaling
  3. CStandardization
  4. DLog Transformation
Show answer & explanation

Correct answer: A. One-Hot Encoding

One-Hot Encoding is the most appropriate technique for nominal categorical features like 'Subscription Type' because it converts them into a numerical format that machine learning models can understand without implying any false ordinal relationships. Each category becomes a new binary feature. Min-Max Scaling, Log Transformation, and Standardization are techniques for numerical features.

Why the other options are wrong

  • B. Min-Max Scaling is used for numerical features to scale them to a specific range, typically 0 to 1.
  • C. Standardization (Z-score normalization) is used for numerical features to transform them to have a mean of 0 and a standard deviation of 1.
  • D. Log Transformation is applied to numerical features, often to reduce skewness or handle outliers.

One-Hot Encoding

A technique used to convert categorical variables into a numerical format that machine learning algorithms can process, where each category is represented as a binary vector.

  • Creates new binary columns for each unique category.
  • Avoids implying ordinal relationships between categories.
  • Suitable for nominal (unordered) categorical data.

Memory trick: Categories become Hot, not just ordered.

More AI/ML and Generative AI Fundamentals questions