AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A data scientist is preparing a large dataset for training a machine learning model. The dataset contains several categorical features with high cardinality, some of which exhibit a skewed distribution where a few categories account for a vast majority of the observations, while many others are rare. The data scientist needs to transform these features to improve model performance and reduce the dimensionality without losing significant information. Which feature engineering technique is most appropriate for this scenario?

  1. ALabel encoding
  2. BOne-hot encoding
  3. CTarget encoding
  4. DBinning numerical features
Show answer & explanation

Correct answer: C. Target encoding

Target encoding replaces each category with the mean of the target variable for that category. This technique is particularly effective for high-cardinality categorical features and can capture the relationship between the feature and the target, reducing dimensionality while preserving information. It inherently handles skewed distributions by focusing on the target's average behavior per category.

Why the other options are wrong

  • A. Label encoding assigns a unique integer to each category, implying an ordinal relationship that might not exist and can mislead models, especially for nominal categories. It does not address high cardinality effectively beyond assigning integers.
  • B. One-hot encoding creates a new binary column for each category, leading to a very high-dimensional sparse matrix for high-cardinality features, which is inefficient and can cause the curse of dimensionality.
  • D. Binning numerical features is used for continuous data, not for categorical features with high cardinality and skewed distributions. It transforms continuous values into discrete bins.

Target Encoding

A feature engineering technique where each category in a categorical feature is replaced by the mean of the target variable for that category.

  • Effective for high-cardinality categorical features.
  • Reduces dimensionality compared to one-hot encoding.
  • Captures the relationship between the feature and the target variable.
  • Can lead to data leakage if not cross-validated properly.

Memory trick: Transforming data wisely, targets reveal the truth for high cards.

More Data Engineering questions