2021
Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
NeurIPS 2021poster
Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovered Maximal Update Parametrization ($\mu$P), many optimal HPs remain stable even as model size changes. This leads to a ne…