← Search

Beomhan Baek

2 accepted papers

2026

Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime

ICLR 2026poster

Adam [Kingma & Ba, 2015] is the de facto optimizer in deep learning, yet its theoretical understanding remains limited. Prior analyses show that Adam favors solutions aligned with $\ell_\infty$-geometry, but these results are restricted to the full-batch regime. In this work, we study the implicit b…

Cited by 1SourceScholar
2025

Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training

NeurIPS 2025poster

As both model and dataset sizes continue to scale rapidly, conventional pretraining strategies with fixed compute budgets—such as cosine learning rate schedules—are increasingly inadequate for large-scale training. Recent alternatives, including warmup-stable-decay (WSD) schedules and weight averagi…

Cited by 0SourceScholar