← Search

Jeremy Welborn

1 accepted papers

2026

Weight Decay may matter more than µP for Learning Rate Transfer in Practice

ICLR 2026poster

Transferring the optimal learning rate from small to large neural networks can enable efficient training at scales where hyperparameter tuning is otherwise prohibitively expensive. To this end, the Maximal Update Parameterization (µP) proposes a learning rate scaling designed to keep the update dyna…

Cited by 0SourcecodeScholar