← Search

Egor Shulgin

9 accepted papers

2026

From Muon to Gluon: Bridging Theory and Practice of LMO-based Optimizers for LLMs

ICML 2026poster

Recent developments in deep learning optimization have brought about radically new algorithms based on the Linear Minimization Oracle (LMO) framework, such as Muon and Scion. After over a decade of Adam's dominance, these LMO-based methods are emerging as viable replacements, offering several practi…

Cited by 0SourceScholar
2026

General Analysis of LMO-based Optimizers: Beyond Bounded Variance

ICML 2026poster

We study a broad family of momentum *Linear Minimization Oracle* (LMO) methods that includes normalized SGD with momentum, sign-based (Adam-like) directions, and Muon (spectral) updates. Our focus is on subsampling regimes where the classical uniformly-bounded-variance model can be fragile even for …

Cited by 0SourceScholar
2025

MAST: model-agnostic sparsified training

ICLR 2025poster

We introduce a novel optimization problem formulation that departs from the conventional way of minimizing machine learning model loss as a black-box function. Unlike traditional formulations, the proposed approach explicitly incorporates an initially pre-trained model and random sketch operators, a…

2021

ADOM: Accelerated Decentralized Optimization Method for Time-Varying Networks

ICML 2021spotlight

We propose ADOM – an accelerated method for smooth and strongly convex decentralized optimization over time-varying networks. ADOM uses a dual oracle, i.e., we assume access to the gradient of the Fenchel conjugate of the individual loss functions. Up to a constant factor, which depends on the netwo…

Cited by 36SourcePDFScholar
2020

Revisiting Stochastic Extragradient

AISTATS 2020poster

We fix a fundamental issue in the stochastic extragradient method by providing a new sampling strategy that is motivated by approximating implicit updates. Since the existing stochastic extragradient algorithm, called Mirror-Prox, of (Juditsky, 2011) diverges on a simple bilinear problem when the do…

Cited by 101SourcePDFScholar
2019

SGD: General Analysis and Improved Rates

ICML 2019oral

We propose a general yet simple theorem describing the convergence of SGD under the arbitrary sampling paradigm. Our theorem describes the convergence of an infinite array of variants of SGD, each of which is associated with a specific probability law governing the data selection rule used to form m…

Cited by 557SourcePDFScholar