2026
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
ICLR 2026poster
We investigate the role of learning rate scheduling in the large-scale pre-training of large language models, focusing on its influence on downstream performance after supervised fine-tuning (SFT). Decay-based learning rate schedulers are widely used to minimize pre-training loss. However, despite t…