AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A data scientist is preparing a dataset for an AI/ML model that will predict the probability of a customer clicking on an advertisement. The dataset contains various numerical features, such as 'Age', 'Income', and 'Number of Past Clicks'. These features have widely different scales and distributions. To prevent features with larger values from dominating the learning process, which data preprocessing technique should be applied?
- AOne-Hot Encoding
- BFeature Scaling
- CBinning
- DText Vectorization
Show answer & explanationAnswer & explanation
Correct answer: B. Feature Scaling
Feature Scaling is the appropriate technique to address the issue of features having widely different scales and distributions. It transforms numerical features to a common scale, preventing features with larger values from disproportionately influencing the model's learning, which is crucial for many machine learning algorithms like gradient descent-based optimizers or distance-based algorithms.
Why the other options are wrong
- A. One-Hot Encoding is used for categorical features, not numerical features with different scales.
- C. Binning converts continuous numerical features into discrete categories or bins, which is different from scaling their values.
- D. Text Vectorization is used to convert text data into numerical representations, irrelevant for numerical features.
Feature Scaling
A data preprocessing technique used to standardize or normalize the range of independent numerical features in a dataset. It ensures that no single feature dominates the learning process due to its magnitude.
- Transforms numerical features to a common scale.
- Prevents features with larger values from dominating model training.
- Includes methods like Standardization (Z-score) and Normalization (Min-Max Scaling).
Memory trick: Scale those numbers, don't let giants dominate!