AWS Certified Machine Learning – SpecialtyModelingHard
A pharmaceutical company is developing a machine learning model to predict the efficacy of new drug compounds. They have a limited dataset of successfully tested compounds. Due to the high cost and time involved in synthesizing and testing new compounds, they need a model that performs well with small datasets and is less prone to overfitting than complex models like deep neural networks. Which type of model is generally preferred for its simplicity and robustness on small datasets?
- ARandom Forest
- BSupport Vector Machine (SVM)
- CGradient Boosting Machine (GBM)
- DDeep Neural Network (DNN)
Show answer & explanationAnswer & explanation
Correct answer: B. Support Vector Machine (SVM)
Support Vector Machines (SVMs) are generally preferred for small datasets because they are less prone to overfitting compared to more complex models. SVMs find an optimal hyperplane that maximizes the margin between classes, effectively generalizing even with limited data points.
Why the other options are wrong
- A. Random Forests are ensemble methods that can work reasonably well on small datasets but might still overfit if the individual trees are too deep. SVMs often provide better generalization guarantees with small data due to their margin-maximization principle.
- C. Gradient Boosting Machines are powerful ensemble methods but can easily overfit on small datasets if not carefully tuned, as they sequentially build complex models.
- D. Deep Neural Networks require large amounts of data to train effectively and are highly prone to overfitting on small datasets due to their high capacity.
Model Simplicity for Small Data
The preference for simpler machine learning models (e.g., SVM, Logistic Regression) over complex ones (e.g., Deep Neural Networks) when dealing with small datasets to prevent overfitting and ensure better generalization.
- Complex models require more data to learn robust patterns.
- Simple models have lower variance, making them less prone to overfitting.
- SVMs are particularly effective due to margin maximization.
Memory trick: Small data, big problem, simple model's the anthem, SVM's the champion, where complexity is banned.