AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A data scientist is tasked with preparing a dataset for an AI/ML model that will predict whether a customer will renew their subscription. The dataset contains a 'Customer ID' column, which uniquely identifies each customer. Which of the following actions should the data scientist take regarding this column during data preparation?
- AUse one-hot encoding to convert 'Customer ID' into numerical features.
- BScale the 'Customer ID' column to normalize its values.
- CRemove the 'Customer ID' column from the dataset.
- DKeep the 'Customer ID' column as is, as it provides unique identification.
Show answer & explanationAnswer & explanation
Correct answer: C. Remove the 'Customer ID' column from the dataset.
Customer IDs are unique identifiers that do not carry predictive power for the model. Including them can lead to overfitting or simply add noise without contributing to the model's ability to generalize patterns from the data.
Why the other options are wrong
- A. One-hot encoding is for categorical features with meaningful categories, not unique identifiers.
- B. Scaling numerical features is appropriate for features with predictive power, not for unique identifiers.
- D. Keeping unique identifiers can cause overfitting as the model might learn to associate a specific outcome with a specific ID, rather than general patterns.
Feature Engineering: Removing IDs
The process of identifying and removing unique identifiers (like customer IDs, transaction IDs) from a dataset before training an AI/ML model.
- Unique identifiers do not contain predictive information.
- Including them can lead to overfitting.
- They add noise and increase dimensionality without benefit.
- Often a crucial step in data preprocessing.
Memory trick: Clean data makes smart models.