AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy

A data scientist is preparing a dataset for an AI/ML model that will predict customer churn. The dataset includes a 'CustomerID' column, which uniquely identifies each customer. What is the most appropriate action for the data scientist to take with this column before training the model?

  1. AUse 'CustomerID' as the target variable for the model.
  2. BRemove the 'CustomerID' column from the dataset.
  3. CConvert 'CustomerID' into a numerical feature using one-hot encoding.
  4. DScale the 'CustomerID' column to a range between 0 and 1.
Show answer & explanation

Correct answer: B. Remove the 'CustomerID' column from the dataset.

Identifier columns like 'CustomerID' are unique to each record and do not contain predictive information for generalized patterns. Including them can lead to overfitting or make the model learn irrelevant correlations, hindering its ability to generalize to new, unseen data. Therefore, they should be removed.

Why the other options are wrong

  • A. The target variable is what the model predicts (e.g., churn), not a unique identifier.
  • C. One-hot encoding is for categorical features, not unique identifiers, and would create too many irrelevant features.
  • D. Scaling is for numerical features that have predictive value, not for unique identifiers.

Removing Identifier Features

The practice of excluding unique identifier columns (like IDs) from a dataset before training a machine learning model.

  • Identifiers are unique to each record.
  • They do not carry predictive power for general patterns.
  • Including them can lead to overfitting.
  • Improves model generalization and performance.

Memory trick: Clean Data: Remove IDs to Avoid Model Confusion

More AI/ML and Generative AI Fundamentals questions