AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A data scientist is preparing a dataset for an AI/ML model that will predict customer churn. The dataset contains a 'Customer ID' column, which is a unique identifier for each customer. Which of the following best describes the appropriate action for this column in the context of model training?
- AUse the 'Customer ID' as a target variable for unsupervised learning.
- BRemove the 'Customer ID' column as it provides no predictive value.
- CApply one-hot encoding to convert the IDs into a numerical format.
- DScale the 'Customer ID' column using normalization or standardization.
Show answer & explanationAnswer & explanation
Correct answer: B. Remove the 'Customer ID' column as it provides no predictive value.
Customer IDs are unique identifiers and typically do not carry predictive information for a model. Including them can lead to overfitting or an inability for the model to generalize. Therefore, they should be removed before training.
Why the other options are wrong
- A. The target variable in this scenario is customer churn, not the customer ID. Customer ID is not a suitable target for unsupervised learning either.
- C. One-hot encoding is used for categorical features with a limited number of unique values, not for unique identifiers like Customer IDs.
- D. Scaling is applied to numerical features that have predictive value, not to unique identifiers that should be removed.
Feature Engineering: Removing IDs
Unique identifiers like customer IDs should generally be removed from datasets before training an AI/ML model because they do not contribute predictive information and can lead to issues like overfitting.
- IDs are unique, not predictive.
- Including IDs can cause overfitting.
- Remove IDs during data preprocessing.
Memory trick: Clean data for smarter decisions.