AWS Certified Machine Learning – SpecialtyModelingHard

A pharmaceutical company is developing a machine learning model to predict the efficacy of new drug compounds based on their chemical structures. They have a limited dataset of successfully tested compounds (around 500 samples), and each sample has a large number of molecular features (over 10,000). The current deep learning model is showing signs of overfitting, with high accuracy on the training set but poor performance on unseen validation data. Which of the following approaches is most appropriate to mitigate overfitting in this scenario?

  1. AIncreasing the complexity of the deep learning model by adding more layers and neurons.
  2. BUsing a more complex ensemble method like a stacked generalization model.
  3. CCollecting significantly more drug compound samples to expand the training dataset.
  4. DApplying Principal Component Analysis (PCA) for dimensionality reduction before training.
Show answer & explanation

Correct answer: D. Applying Principal Component Analysis (PCA) for dimensionality reduction before training.

With a limited number of samples (500) and a very high number of features (10,000), the model is likely suffering from the 'curse of dimensionality,' leading to overfitting. PCA can effectively reduce the number of features while retaining most of the variance, making the problem more manageable for the limited data and mitigating overfitting.

Why the other options are wrong

  • A. Increasing model complexity would exacerbate overfitting, not mitigate it, given the limited data.
  • B. More complex ensemble methods can also overfit if the base models are already overfitting or if the dimensionality issue isn't addressed first.
  • C. Collecting more data is ideal but often not feasible in pharmaceutical research due to high costs and time; the question asks for the most appropriate *approach* given the current scenario.

Dimensionality Reduction for Overfitting

Techniques like PCA that reduce the number of features in a dataset, which is particularly useful for mitigating overfitting when the number of features significantly exceeds the number of samples.

  • Reduces computational cost.
  • Helps combat the curse of dimensionality.
  • Can improve model generalization by removing noise.

Memory trick: Reduce Dimensions, Retain Wisdom, Avoid Over-Fit.

More Modeling questions