← Search

Anton Rodomanov

13 accepted papers

2026

Composite Optimization with Error Feedback: the Dual Averaging Approach

ICLR 2026poster

Communication efficiency is a central challenge in distributed machine learning training, and message compression is a widely used solution. However, standard Error Feedback (EF) methods (Seide et al., 2014), though effective for smooth unconstrained optimization with compression (Karimireddy et al.…

Cited by 0SourceScholar
2026

DADA: Dual Averaging with Distance Adaptation

ICLR 2026poster

We present a novel parameter-free universal gradient method for solving convex optimization problems. Our algorithm—Dual Averaging with Distance Adaptation (DADA)–is based on the classical scheme of dual averaging and dynamically adjusts its coefficients based on the observed gradients and the dista…

Cited by 0SourceScholar
2025

Decoupled SGDA for Games with Intermittent Strategy Communication

ICML 2025poster

We introduce *Decoupled SGDA*, a novel adaptation of Stochastic Gradient Descent Ascent (SGDA) tailored for multiplayer games with intermittent strategy communication. Unlike prior methods, Decoupled SGDA enables players to update strategies locally using outdated opponent strategies, significantly…

Cited by 0SourcePDFScholar
2025

Exploiting Similarity for Computation and Communication-Efficient Decentralized Optimization

ICML 2025poster

Reducing communication complexity is critical for efficient decentralized optimization. The proximal decentralized optimization (PDO) framework is particularly appealing, as methods within this framework can exploit functional similarity among nodes to reduce communication rounds. Specifically, when…

Cited by 0SourcePDFScholar
2025

Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods

ICLR 2025poster

We study gradient methods for optimizing $(L_0, L_1)$-smooth functions, a class that generalizes Lipschitz-smooth functions and has gained attention for its relevance in machine learning. We provide new insights into the structure of this function class and develop a principled framework for analyzi…

Cited by 0SourcePDFScholar
2024

Federated Optimization with Doubly Regularized Drift Correction

ICML 2024poster

Federated learning is a distributed optimization paradigm that allows training machine learning models across decentralized devices while keeping the data localized. The standard method, FedAvg, suffers from client drift which can hamper performance and increase communication costs over centralized…

Cited by 18SourcePDFScholar
2024

Stabilized Proximal-Point Methods for Federated Optimization

NeurIPS 2024spotlight

In developing efficient optimization algorithms, it is crucial to account for communication constraints—a significant challenge in modern Federated Learning. The best-known communication complexity among non-accelerated algorithms is achieved by DANE, a distributed proximal-point algorith…

2024

Universal Gradient Methods for Stochastic Convex Optimization

ICML 2024poster

We develop universal gradient methods for Stochastic Convex Optimization (SCO). Our algorithms automatically adapt not only to the oracle's noise but also to the Hölder smoothness of the objective function without a priori knowledge of the particular setting. The key ingredient is a novel strategy f…

Cited by 2SourcePDFScholar
2024

Universality of AdaGrad Stepsizes for Stochastic Optimization: Inexact Oracle, Acceleration and Variance Reduction

NeurIPS 2024poster

We present adaptive gradient methods (both basic and accelerated) for solving convex composite optimization problems in which the main part is approximately smooth (a.k.a. $(\delta, L)$-smooth) and can be accessed only via a (potentially biased) stochastic gradient oracle. This setting covers many i…

Cited by 4SourcePDFScholar
2016

A Superlinearly-Convergent Proximal Newton-type Method for the Optimization of Finite Sums

ICML 2016poster

We consider the problem of minimizing the strongly convex sum of a finite number of convex functions. Standard algorithms for solving this problem in the class of incremental/stochastic methods have at most a linear convergence rate. We propose a new incremental method whose convergence rate is supe…

Cited by 51SourcePDFScholar