AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium

A data engineer is preparing a dataset for an AI/ML model that will predict whether a customer will renew their subscription. The dataset contains a unique 'CustomerID' for each record. If this feature is directly included in the model training, it could lead to the model memorizing individual customer details rather than learning general patterns, potentially resulting in poor generalization. Which feature engineering technique should be applied to address this issue?

  1. AFeature Scaling
  2. BPolynomial Feature Generation
  3. CRemoving Identifier Features
  4. DOne-Hot Encoding
Show answer & explanation

Correct answer: C. Removing Identifier Features

Identifier features like 'CustomerID' are unique to each record and do not carry predictive power for general patterns. Including them directly can cause the model to overfit by memorizing specific IDs. Therefore, the appropriate feature engineering technique is to remove such identifier features from the dataset before training.

Why the other options are wrong

  • A. Feature Scaling normalizes numerical features, but 'CustomerID' is an identifier, not a meaningful numerical feature to scale.
  • B. Polynomial Feature Generation creates new features by raising existing numerical features to a power, which is not applicable or beneficial for unique identifiers.
  • D. One-Hot Encoding a unique identifier would create a new binary column for every customer, leading to an extremely sparse, high-dimensional dataset and severe overfitting.

Removing Identifier Features

A feature engineering technique where unique identification columns (e.g., IDs, serial numbers) are removed from the dataset before model training, as they typically do not contribute to learning generalizable patterns and can cause overfitting.

  • Prevents models from memorizing specific instances.
  • Reduces dimensionality and potential for overfitting.
  • Applies to features like CustomerID, OrderID, SSN.

Memory trick: IDs are unique, not general, so remove them.

More AI/ML and Generative AI Fundamentals questions