← Search

Ilyas Fatkhullin

12 accepted papers

2025

Safe-EF: Error Feedback for Non-smooth Constrained Optimization

ICML 2025poster

Federated learning faces severe communication bottlenecks due to the high dimensionality of model updates. Communication compression with contractive compressors (e.g., Top-$K$) is often preferable in practice but can degrade performance without proper handling. Error feedback (EF) mitigates such is…

2025

Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits

NeurIPS 2025poster

Heavy-tailed noise is pervasive in modern machine learning applications, arising from data heterogeneity, outliers, and non-stationary stochastic environments. While second-order methods can significantly accelerate convergence in light-tailed or bounded-noise settings, such algorithms are often bri…

Cited by 0SourceScholar
2023

Reinforcement Learning with General Utilities: Simpler Variance Reduction and Large State-Action Space

ICML 2023poster

We consider the reinforcement learning (RL) problem with general utilities which consists in maximizing a function of the state-action occupancy measure. Beyond the standard cumulative reward RL setting, this problem includes as particular cases constrained RL, pure exploration and learning from dem…

Cited by 19SourcePDFScholar
2023

Stochastic Policy Gradient Methods: Improved Sample Complexity for Fisher-non-degenerate Policies

ICML 2023poster

Recently, the impressive empirical success of policy gradient (PG) methods has catalyzed the development of their theoretical foundations. Despite the huge efforts directed at the design of efficient stochastic PG-type algorithms, the understanding of their convergence to a globally optimal policy i…

Cited by 50SourcePDFScholar
2023

Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive Methods

NeurIPS 2023poster

The classical analysis of Stochastic Gradient Descent (SGD) with polynomially decaying stepsize $\eta_t = \eta/\sqrt{t}$ relies on well-tuned $\eta$ depending on problem parameters such as Lipschitz smoothness constant, which is often unknown in practice. In this work, we prove that SGD with arbitr…

Cited by 30SourcePDFScholar
2022

3PC: Three Point Compressors for Communication-Efficient Distributed Training and a Better Theory for Lazy Aggregation

ICML 2022spotlight

We propose and study a new class of gradient compressors for communication-efficient training—three point compressors (3PC)—as well as efficient distributed nonconvex optimization algorithms that can take advantage of them. Unlike most established approaches, which rely on a static compressor choice…

Cited by 36SourcePDFScholar
2022

Sharp Analysis of Stochastic Optimization under Global Kurdyka-Lojasiewicz Inequality

NeurIPS 2022accept

We study the complexity of finding the global solution to stochastic nonconvex optimization when the objective function satisfies global Kurdyka-{\L}ojasiewicz (KL) inequality and the queries from stochastic gradient oracles satisfy mild expected smoothness assumption. We first introduce a general…

Cited by 32SourcePDFScholar
2021

EF21: A New, Simpler, Theoretically Better, and Practically Faster Error Feedback

NeurIPS 2021oral

Error feedback (EF), also known as error compensation, is an immensely popular convergence stabilization mechanism in the context of distributed training of supervised machine learning models enhanced by the use of contractive communication compression mechanisms, such as Top-$k$. First proposed by…

Cited by 177SourcePDFScholar