← Search

Shaheer Muhammad

1 accepted papers

2026

Scaling with Collapse: Efficient and Predictable Training of LLM Families

ICLR 2026poster

Effective LLM training relies on *consistency*, meaning that key quantities—such as final losses and optimal hyperparameters—scale predictably across model sizes. Qiu et al. (2025) recently showed that this consistency extends beyond scalars: whole training loss curves can *collapse* onto a universa…

Cited by 0SourceScholar