2026
Weight Decay may matter more than µP for Learning Rate Transfer in Practice
ICLR 2026poster
Transferring the optimal learning rate from small to large neural networks can enable efficient training at scales where hyperparameter tuning is otherwise prohibitively expensive. To this end, the Maximal Update Parameterization (µP) proposes a learning rate scaling designed to keep the update dyna…