AWS Certified Machine Learning – SpecialtyModelingMedium

A data science team is developing a machine learning model to predict the likelihood of a rare medical condition. The dataset is extremely small due to the rarity of the condition, and the team needs to ensure the model generalizes well to unseen patient data. They are considering different model architectures. Which approach is generally best suited for this scenario?

  1. AEmploying a complex ensemble method like a deep stacking model.
  2. BLeveraging a large pre-trained transformer model and fine-tuning it with the small dataset.
  3. CTraining a very deep neural network with millions of parameters.
  4. DUsing a simple, interpretable model such as Logistic Regression or a Decision Tree.
Show answer & explanation

Correct answer: D. Using a simple, interpretable model such as Logistic Regression or a Decision Tree.

With extremely small datasets, complex models (like deep neural networks, deep ensembles, or large pre-trained transformers) are highly prone to overfitting because they have too many parameters relative to the available data. Simple, interpretable models like Logistic Regression or Decision Trees have fewer parameters and are less likely to overfit, making them better suited for scenarios with limited data to ensure better generalization.

Why the other options are wrong

  • A. Complex ensemble methods, especially deep stacking, also tend to have high capacity and are susceptible to overfitting on small datasets.
  • B. Large pre-trained transformer models, even with fine-tuning, still have a massive number of parameters and can easily overfit extremely small, domain-specific datasets without careful regularization or extensive data augmentation.
  • C. Very deep neural networks require large amounts of data to train effectively and are highly prone to overfitting with small datasets.

Model Simplicity for Small Data

For datasets with limited observations, simpler machine learning models are generally preferred over complex ones to reduce the risk of overfitting and improve generalization.

  • Complex models with many parameters require large amounts of data to learn effectively.
  • Simple models (e.g., linear models, shallow trees) have lower capacity and are less prone to overfitting.
  • Prioritize interpretability and robustness when data is scarce.

Memory trick: When data is 'tiny', keep your model 'simple' like a 'toy', or it will 'overfit' and give you 'noise'.

More Modeling questions