← Search

Yelin He

4 accepted papers

2026

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD

ICLR 2026poster

Transformers have become the de facto backbone of modern deep learning, yet their training typically demands an advanced optimizer with adaptive learning rate like AdamW, rather than a momentum SGDW (mSGDW). Previous works show that it is mainly due to a heavy-tailed distribution of the gradients. I…

Cited by 0SourceScholar
2025

Taming Transformer Without Using Learning Rate Warmup

ICLR 2025poster

Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and an obviously lower learning rate, is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Tran…

Cited by 0SourcePDFScholar