← Search

Elnur Gasanov

7 accepted papers

2024

Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants

ICLR 2024poster

Error feedback (EF) is a highly popular and immensely effective mechanism for fixing convergence issues which arise in distributed training methods (such as distributed GD or SGD) when these are enhanced with greedy communication compression techniques such as Top-k. While EF was proposed almost a d…

Cited by 3SourcePDFScholar
2024

Understanding Progressive Training Through the Framework of Randomized Coordinate Descent

AISTATS 2024poster

We propose a Randomized Progressive Training algorithm (RPT)—a stochastic proxy for the well-known Progressive Training method (PT) (Karras et al., 2017). Originally designed to train GANs (Goodfellow et al., 2014), PT was proposed as a heuristic, with no convergence analysis even for the simplest o…

Cited by 4SourcePDFScholar
2022

3PC: Three Point Compressors for Communication-Efficient Distributed Training and a Better Theory for Lazy Aggregation

ICML 2022spotlight

We propose and study a new class of gradient compressors for communication-efficient training—three point compressors (3PC)—as well as efficient distributed nonconvex optimization algorithms that can take advantage of them. Unlike most established approaches, which rely on a static compressor choice…

Cited by 36SourcePDFScholar
2022

FLIX: A Simple and Communication-Efficient Alternative to Local Methods in Federated Learning

AISTATS 2022poster

Federated Learning (FL) is an increasingly popular machine learning paradigm in which multiple nodes try to collaboratively learn under privacy, communication and multiple heterogeneity constraints. A persistent problem in federated learning is that it is not clear what the optimization objective sh…

2021

Lower Bounds and Optimal Algorithms for Smooth and Strongly Convex Decentralized Optimization Over Time-Varying Networks

NeurIPS 2021poster

We consider the task of minimizing the sum of smooth and strongly convex functions stored in a decentralized manner across the nodes of a communication network whose links are allowed to change in time. We solve two fundamental problems for this task. First, we establish {\em the first lower bounds}…

Cited by 47SourcePDFScholar
2020

From Local SGD to Local Fixed-Point Methods for Federated Learning

ICML 2020poster

Most algorithms for solving optimization problems or finding saddle points of convex-concave functions are fixed-point algorithms. In this work we consider the generic problem of finding a fixed point of an average of operators, or an approximation thereof, in a distributed setting. Our work is moti…

Cited by 156SourcePDFScholar
2018

Stochastic Spectral and Conjugate Descent Methods

NeurIPS 2018poster

The state-of-the-art methods for solving optimization problems in big dimensions are variants of randomized coordinate descent (RCD). In this paper we introduce a fundamentally new type of acceleration strategy for RCD based on the augmentation of the set of coordinate directions by a few spectral o…

Cited by 15SourcePDFScholar