AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A data scientist is preparing a dataset for an AI/ML model that will predict customer churn. One of the features in the dataset is 'Customer_ID', which is a unique alphanumeric identifier for each customer. What is the most appropriate action for the data scientist to take regarding this feature before training the model?
- AEncode 'Customer_ID' using one-hot encoding to convert it into a numerical format.
- BScale 'Customer_ID' using standardization to ensure it has a mean of 0 and a standard deviation of 1.
- CTransform 'Customer_ID' into a categorical feature by grouping similar IDs together.
- DRemove 'Customer_ID' from the dataset as it is unlikely to contribute to the model's predictive power.
Show answer & explanationAnswer & explanation
Correct answer: D. Remove 'Customer_ID' from the dataset as it is unlikely to contribute to the model's predictive power.
Unique identifiers like 'Customer_ID' typically do not provide predictive power for the target variable and can introduce noise or even lead to overfitting if improperly handled. Removing them is the standard practice.
Why the other options are wrong
- A. One-hot encoding unique identifiers would create an extremely sparse matrix with no predictive value and increase dimensionality significantly, which is inefficient.
- B. Scaling numerical features is appropriate, but Customer_ID is an identifier, not a meaningful numerical feature that would benefit from scaling.
- C. Grouping unique IDs would be arbitrary and would not create meaningful categorical features relevant to customer churn prediction.
Removing Identifier Features
The process of eliminating unique identifiers (e.g., IDs, timestamps) from a dataset before training an AI/ML model because they typically do not contain predictive information.
- Identifiers are unique to each data point.
- They usually don't generalize across data for prediction.
- Can lead to overfitting or high dimensionality if used as features.
Memory trick: Engineers Clean Data to Predict.