AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy

A data scientist is preparing a dataset for an AI/ML model that will predict customer churn. The dataset contains a 'Customer ID' column, which is a unique identifier for each customer. Which of the following best describes the appropriate action for this column in the context of model training?

  1. AUse the 'Customer ID' as a target variable for unsupervised learning.
  2. BRemove the 'Customer ID' column as it provides no predictive value.
  3. CApply one-hot encoding to convert the IDs into a numerical format.
  4. DScale the 'Customer ID' column using normalization or standardization.
Show answer & explanation

Correct answer: B. Remove the 'Customer ID' column as it provides no predictive value.

Customer IDs are unique identifiers and typically do not carry predictive information for a model. Including them can lead to overfitting or an inability for the model to generalize. Therefore, they should be removed before training.

Why the other options are wrong

  • A. The target variable in this scenario is customer churn, not the customer ID. Customer ID is not a suitable target for unsupervised learning either.
  • C. One-hot encoding is used for categorical features with a limited number of unique values, not for unique identifiers like Customer IDs.
  • D. Scaling is applied to numerical features that have predictive value, not to unique identifiers that should be removed.

Feature Engineering: Removing IDs

Unique identifiers like customer IDs should generally be removed from datasets before training an AI/ML model because they do not contribute predictive information and can lead to issues like overfitting.

  • IDs are unique, not predictive.
  • Including IDs can cause overfitting.
  • Remove IDs during data preprocessing.

Memory trick: Clean data for smarter decisions.

More AI/ML and Generative AI Fundamentals questions