Microsoft Azure AI Fundamentals (AI-900)Describe fundamental principles of machine learning on AzureEasy
A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset includes a feature called 'Customer_ID', which is a unique identifier for each customer. This feature has no inherent numerical meaning or order. What is the most appropriate action for this feature before training the model?
- ARemove 'Customer_ID' from the dataset.
- BConvert 'Customer_ID' to a one-hot encoded vector.
- CApply a logarithmic transformation to 'Customer_ID'.
- DNormalize 'Customer_ID' to a range between 0 and 1.
Show answer & explanationAnswer & explanation
Correct answer: A. Remove 'Customer_ID' from the dataset.
Unique identifiers like 'Customer_ID' do not carry predictive power for customer churn and can introduce noise or lead to overfitting if included. Therefore, they should be removed from the dataset.
Why the other options are wrong
- B. One-hot encoding is for categorical features with meaningful categories, not unique identifiers.
- C. Logarithmic transformation is for skewed numerical data, which 'Customer_ID' is not in a meaningful way.
- D. Normalizing a unique identifier without numerical meaning is inappropriate and will not make it useful for prediction.
Feature Selection
The process of selecting a subset of relevant features (variables, predictors) for use in model construction.
- Improves model performance by reducing overfitting.
- Reduces training time and computational cost.
- Enhances model interpretability.
Memory trick: Clean Data, Clear Decisions