← Search

Sarit Khirirat

8 accepted papers

2025

Error Feedback under $(L_0,L_1)$-Smoothness: Normalization and Momentum

NeurIPS 2025poster

We provide the first proof of convergence for normalized error feedback algorithms across a wide range of machine learning problems. Despite their popularity and efficiency in training deep neural networks, traditional analyses of error feedback algorithms rely on the smoothness assumption that doe…

Cited by 0SourceScholar
2022

Eco-Fedsplit: Federated Learning with Error-Compensated Compression

ICASSP 2022accepted

Federated learning is an emerging framework for collaborative machine-learning on devices which do not want to share local data. State-of-the art methods in federated learning reduce the communication frequency, but are not guaranteed to converge to the optimal model parameters. These methods also e…

Cited by 0SourceScholar
2021

A Flexible Framework for Communication-Efficient Machine Learning

AAAI 2021technical

With the increasing scale of machine learning tasks, it has become essential to reduce the communication between computing nodes. Early work on gradient compression focused on the bottleneck between CPUs and GPUs, but communication-efficiency is now needed in a variety of different system architec…

Cited by 17SourcePDFScholar
2021

Improved Step-Size Schedules for Noisy Gradient Methods

ICASSP 2021accepted

Noise is inherited in many optimization methods such as stochastic gradient methods, zeroth-order methods and compressed gradient methods. For such methods to converge toward a global optimum, it is intuitive to use large step-sizes in the initial iterations when the noise is typically small compare…

Cited by 0SourceScholar
2019

Convergence Bounds for Compressed Gradient Methods with Memory Based Error Compensation

ICASSP 2019accepted

The veritable scale of modern data necessitates information compression in parallel/distributed big-data optimization. Compression schemes using memory-based error compensation have displayed superior performance in practice, however, to date there are no theoretical explanations for these observed…

Cited by 0SourceScholar
2018

The Convergence of Sparsified Gradient Methods

NeurIPS 2018poster

Distributed training of massive machine learning models, in particular deep neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace. Several families of communication-reduction methods, such as quantization, large-batch methods, and gradient sparsification, have been proposed.…

Cited by 640SourcePDFScholar