AWS Certified Machine Learning – SpecialtyModelingMedium
A research team is developing a machine learning model to predict the progression of a rare disease. They have a very small dataset of patient records, which includes various clinical measurements and the disease progression outcome. Due to the rarity of the disease, collecting more data is extremely difficult. The team's initial attempts with complex models show high variance and poor generalization. Which modeling strategy is most appropriate for this scenario?
- AUtilize transfer learning by fine-tuning a pre-trained model from a related, large dataset.
- BUse a simple, interpretable model like Logistic Regression or a Decision Tree.
- CTrain a very deep neural network with millions of parameters.
- DApply extensive data augmentation techniques to synthetically increase the dataset size.
Show answer & explanationAnswer & explanation
Correct answer: B. Use a simple, interpretable model like Logistic Regression or a Decision Tree.
With a very small dataset, complex models are highly prone to overfitting (high variance) and will not generalize well. Simple, interpretable models have fewer parameters, making them less likely to overfit sparse data and more robust in such scenarios.
Why the other options are wrong
- A. Transfer learning requires a large, related pre-trained model, which might not exist or be suitable for rare tabular disease data, and fine-tuning with very little data can still lead to overfitting.
- C. Very deep neural networks require large datasets to train effectively and would almost certainly overfit a very small dataset.
- D. While data augmentation can help, it is often more effective for image/audio data and might not fully address the fundamental issues of a tiny, complex, tabular dataset. Moreover, it can introduce synthetic biases if not done carefully.
Model Simplicity for Small Data
When training data is scarce, simpler models are preferred to avoid overfitting and ensure better generalization.
- Complex models have high capacity and easily overfit small datasets.
- Simple models (e.g., linear models, shallow trees) have fewer parameters.
- Reduces variance, leading to more robust predictions on unseen data.
Memory trick: Small data, simple model.