← Search

Yufei Gu

5 accepted papers

2026

Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better

ICLR 2026poster

As Large Language Models (LLMs) achieve remarkable empirical success through scaling model and data size, pretraining has become increasingly critical yet computationally prohibitive, hindering rapid development. Despite the availability of numerous pretrained LLMs developed at significant computati…

Cited by 0SourceScholar
2025

Investigating the Overlooked Hessian Structure: From CNNs to LLMs

ICML 2025poster

It is well-known that the Hessian of deep loss landscape matters to optimization and generalization of deep learning. Previous studies reported a rough Hessian structure in deep learning, which consists of two components, a small number of large eigenvalues and a large number of nearly-zero eigenval…

Cited by 0SourcePDFScholar
2024

Unraveling the Enigma of Double Descent: An In-depth Analysis through the Lens of Learned Feature Space

ICLR 2024poster

Double descent presents a counter-intuitive aspect within the machine learning domain, and researchers have observed its manifestation in various models and tasks. While some theoretical explanations have been proposed for this phenomenon in specific contexts, an accepted theory for its occurring me…