AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsEasy
A data scientist is preparing a dataset for an AI/ML model that will predict the probability of a customer defaulting on a loan. The dataset contains various numerical features such as 'LoanAmount' (ranging from $1,000 to $1,000,000) and 'CreditScore' (ranging from 300 to 850). What data preprocessing technique should be applied to these features to ensure that features with larger numerical ranges do not disproportionately influence the model?
- AOne-hot encoding
- BRemoving outliers
- CImputing missing values
- DFeature scaling
Show answer & explanationAnswer & explanation
Correct answer: D. Feature scaling
Feature scaling is essential when features have vastly different numerical ranges. Algorithms that calculate distances between data points (like K-Nearest Neighbors, Support Vector Machines, or neural networks) can be heavily biased by features with larger scales. Scaling ensures all features contribute proportionally to the model's learning process, preventing features with larger values from dominating.
Why the other options are wrong
- A. One-hot encoding is used for categorical features, not numerical features with different ranges.
- B. Removing outliers addresses extreme values but doesn't solve the issue of disproportionate influence due to scale differences.
- C. Imputing missing values handles gaps in the data, which is a different preprocessing concern.
Feature Scaling
A data preprocessing technique used to normalize the range of independent variables or features of data.
- Prevents features with larger numerical values from dominating the model's learning.
- Crucial for distance-based algorithms (e.g., SVM, KNN, neural networks).
- Common methods include Min-Max Scaling (Normalization) and Standardization (Z-score normalization).
- Ensures all features contribute equally to the model's performance.
Memory trick: Scale the Features, Don't Let Big Numbers Bully!