AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium
A data scientist is investigating a trained AI model that predicts loan default risk. They discover that the model consistently predicts a higher default risk for applicants from a particular zip code, even when controlling for other financial factors. This leads to a disproportionately higher rate of loan rejections for residents of that zip code. What type of issue does this scenario represent?
- AData Leakage
- BAlgorithmic Bias
- CUnderfitting
- DOverfitting
Show answer & explanationAnswer & explanation
Correct answer: B. Algorithmic Bias
This scenario describes Algorithmic Bias, where the AI model produces systematically unfair or discriminatory outcomes for a particular group (residents of a specific zip code) due to skewed training data or flawed model design. The model is not just making errors, but exhibiting a pattern of disadvantage against a group, even if unintended.
Why the other options are wrong
- A. Data Leakage occurs when information from outside the training dataset is used to create the model, leading to overly optimistic performance estimates.
- C. Underfitting means the model is too simple to capture the underlying patterns in the data, performing poorly on both training and test data.
- D. Overfitting means the model performs well on training data but poorly on unseen data due to memorizing noise.
Algorithmic Bias
Systematic and unfair discrimination by an AI/ML algorithm against certain individuals or groups, often stemming from biases present in the training data, model design, or evaluation metrics.
- Can lead to discriminatory outcomes (e.g., loan rejections, hiring).
- Often unintentional, but has significant societal impact.
- Requires careful data auditing, model explainability, and fairness metrics to mitigate.
Memory trick: Biased algorithms are unfair, not just wrong.