AWS Certified Machine Learning – SpecialtyModelingMedium

A data science team is developing a machine learning model to predict the likelihood of a rare medical condition based on patient data. The dataset available is very small, containing only 200 patient records, with a highly imbalanced class distribution for the rare condition. The team is concerned about overfitting and generalization performance due to the limited data. Which modeling approach is most appropriate for this scenario?

  1. AEmploying a complex ensemble method like XGBoost with default hyperparameters.
  2. BImplementing a large Transformer model pre-trained on a massive text corpus, then fine-tuning it.
  3. CTraining a deep neural network with millions of parameters and extensive data augmentation.
  4. DUtilizing a simple, interpretable model such as Logistic Regression or a Decision Tree with regularization.
Show answer & explanation

Correct answer: D. Utilizing a simple, interpretable model such as Logistic Regression or a Decision Tree with regularization.

With a very small dataset and concerns about overfitting, simple, interpretable models with regularization are preferred. These models have fewer parameters, making them less prone to memorizing the training data and better at generalizing to unseen data when data is scarce. Complex models require large amounts of data to learn effectively.

Why the other options are wrong

  • A. Complex ensemble methods like XGBoost can still overfit small datasets if not carefully tuned, and their complexity might be unnecessary.
  • B. Transformer models are designed for large text data and would be severely overparameterized and ineffective for a small tabular medical dataset.
  • C. Deep neural networks with millions of parameters are highly prone to overfitting on small datasets.

Model Simplicity for Small Data

When working with small datasets, simpler machine learning models are generally preferred over complex ones to prevent overfitting and ensure better generalization.

  • Complex models have many parameters, requiring extensive data to learn robust patterns.
  • Simple models have fewer parameters, reducing the risk of memorizing noise in small datasets.
  • Regularization techniques can further help simple models generalize better.

Memory trick: Small data, small model, big generalization.

More Modeling questions