A healthcare startup is developing a machine learning model to diagnose a rare disease from patient medical records. The dataset is extremely small, with only 100 positive cases and 1000 negative cases. Due to the rarity of the disease and the high cost of data collection, acquiring more data is not feasible. The team initially tried a complex deep learning model, but it showed severe overfitting. Which modeling approach is generally most suitable for such a scenario (very small dataset, limited positive samples, high cost of data acquisition)?
- AImplementing a Generative Adversarial Network (GAN) to generate synthetic data.
- BTraining a very deep convolutional neural network.
- CEmploying a simple, interpretable model such as Logistic Regression or a Decision Tree.
- DUtilizing a complex ensemble method like XGBoost with many estimators.
Show answer & explanationAnswer & explanation
Correct answer: C. Employing a simple, interpretable model such as Logistic Regression or a Decision Tree.
With a very small dataset, especially with limited positive samples, complex models like deep neural networks or large ensembles are highly prone to overfitting. Simple, interpretable models such as Logistic Regression or Decision Trees have fewer parameters, making them less likely to overfit sparse data and often provide better generalization on small datasets. They also offer interpretability, which is crucial in healthcare.
Why the other options are wrong
- A. GANs for synthetic data generation are complex to train and require substantial real data themselves to produce high-quality, diverse synthetic samples, making them unsuitable for an already extremely small dataset.
- B. Very deep CNNs require large amounts of data to train effectively and would severely overfit a small dataset.
- D. XGBoost with many estimators, while powerful, can also overfit if not carefully tuned, especially with such limited data.
Model Simplicity for Small Data
The principle of selecting simpler machine learning models (e.g., linear models, basic decision trees) when dealing with very small datasets to prevent overfitting and improve generalization, often at the expense of capturing highly complex patterns.
- Reduces the risk of overfitting.
- Requires fewer parameters to learn.
- Often more interpretable.
- Can outperform complex models on sparse data.
Memory trick: Keep It Simple, Stupid: Small Data, Simple Model.