AWS Certified Machine Learning – SpecialtyModelingMedium

A research team is training a deep neural network for medical image classification. They observe that the model achieves very high accuracy on the training data but performs poorly on new, unseen data, indicating overfitting. The team has already tried increasing the dataset size and adding dropout layers. To further mitigate overfitting and improve generalization without significantly increasing training time, which hyperparameter adjustment for the Adam optimizer should they consider, and why?

  1. AIncreasing the learning rate (e.g., from 0.001 to 0.01) to speed up convergence.
  2. BIncreasing the `epsilon` parameter to prevent division by zero during optimization.
  3. CIncreasing the `weight_decay` (L2 regularization) parameter to penalize large weights.
  4. DDecreasing the batch size (e.g., from 128 to 32) to introduce more noise in gradients.
Show answer & explanation

Correct answer: C. Increasing the `weight_decay` (L2 regularization) parameter to penalize large weights.

Overfitting implies the model has learned too much from the training data, including noise. Weight decay (L2 regularization) is a common technique to mitigate overfitting by adding a penalty to the loss function proportional to the square of the magnitude of the weights. This encourages the model to use smaller weights, leading to a simpler model and better generalization. Adam optimizer supports `weight_decay` as a hyperparameter.

Why the other options are wrong

  • A. Increasing the learning rate might make the model converge faster but could also lead to instability or divergence, and does not directly address overfitting; it might even worsen it by overshooting optimal weights.
  • B. Increasing `epsilon` in Adam is a numerical stability parameter to prevent division by zero and has no direct impact on mitigating overfitting.
  • D. Decreasing the batch size can add more noise to gradients, which might act as a mild regularizer, but its primary effect is on stability and speed of convergence, not as direct an overfitting mitigation as weight decay.

Weight Decay (L2 Regularization)

Weight decay, or L2 regularization, is a technique used to prevent overfitting by adding a penalty term to the loss function that is proportional to the square of the magnitude of the model's weights.

  • Encourages smaller weights, leading to simpler models.
  • Reduces the model's sensitivity to small changes in input.
  • Helps improve generalization to unseen data.
  • A hyperparameter that needs to be tuned (e.g., in optimizers like Adam, SGD).

Memory trick: To stop 'OVER-LEARNING', make your model 'SIMPLER, NOISIER, or REGULATED'.

More Modeling questions