AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A machine learning team is working with a dataset that contains various categorical features, some with high cardinality (e.g., product IDs, user IDs). They observe that one-hot encoding these features leads to an extremely sparse dataset with a very high number of dimensions, negatively impacting model training time and memory usage. They want to reduce the dimensionality while still capturing useful information from these categorical features. Which feature engineering technique should they employ?

  1. AFrequency Encoding
  2. BLabel Encoding
  3. CBinning
  4. DPrincipal Component Analysis (PCA)
Show answer & explanation

Correct answer: A. Frequency Encoding

Frequency encoding replaces each category with the count or frequency of its occurrence in the dataset. This technique helps reduce dimensionality for high-cardinality categorical features while still providing numerical information that can be useful for models, especially when the frequency of a category is correlated with the target variable.

Why the other options are wrong

  • B. Label encoding assigns a unique integer to each category. While it reduces dimensionality, it introduces an arbitrary ordinal relationship that can mislead models if no such relationship exists.
  • C. Binning is typically used for numerical features to group continuous values into discrete bins, not directly for reducing cardinality of existing categorical features.
  • D. PCA is a dimensionality reduction technique primarily for numerical features. Applying it directly to one-hot encoded sparse data might not be optimal and can be computationally expensive for very high dimensions before reduction.

Frequency Encoding

A categorical feature encoding technique where each category is replaced by the frequency or count of its occurrence in the dataset. It's useful for high-cardinality features to reduce dimensionality.

  • Reduces dimensionality for high-cardinality categorical features
  • Preserves information about category prevalence
  • Does not introduce arbitrary ordinality
  • Can be effective if frequency correlates with the target

Memory trick: When too many categories make a mess, frequency encoding is your best address.

More Data Engineering questions