AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium
A data engineer is preparing a dataset for an AI/ML model that will predict the probability of a customer clicking on an advertisement. The dataset contains numerical features with vastly different scales, such as 'age' (ranging from 18-90) and 'income' (ranging from $20,000-$500,000). To prevent features with larger numerical values from disproportionately influencing the model's learning process, what preprocessing technique should be applied?
- AData augmentation
- BOne-hot encoding
- CDimensionality reduction
- DFeature scaling
Show answer & explanationAnswer & explanation
Correct answer: D. Feature scaling
Feature scaling (e.g., normalization or standardization) is essential when features have different ranges. It ensures that all features contribute proportionally to the model's learning, preventing features with larger values from dominating the optimization process.
Why the other options are wrong
- A. Data augmentation is used to increase the size and diversity of a training dataset, typically for image or text data, and is not relevant to scaling numerical features.
- B. One-hot encoding is for converting categorical data into a numerical format, not for handling numerical features with different scales.
- C. Dimensionality reduction (e.g., PCA) reduces the number of features, which can help with model complexity but doesn't directly address the issue of different scales among existing features.
Feature Scaling
Feature scaling is a data preprocessing technique used to standardize or normalize the range of independent numerical features in a dataset. This prevents features with larger values from dominating the model's learning process and helps optimization algorithms converge faster.
- Standardizes/normalizes numerical features.
- Prevents dominance by large values.
- Helps model convergence.
Memory trick: Clean and shape data for superior models.