← Search

Farzin Haddadpour

6 accepted papers

2023

Beyond Lipschitz: Sharp Generalization and Excess Risk Bounds for Full-Batch GD

ICLR 2023poster

We provide sharp path-dependent generalization and excess risk guarantees for the full-batch Gradient Descent (GD) algorithm on smooth losses (possibly non-Lipschitz, possibly nonconvex). At the heart of our analysis is an upper bound on the generalization error, which implies that average output st…

Cited by 19SourcePDFScholar
2022

Black-Box Generalization: Stability of Zeroth-Order Learning

NeurIPS 2022accept

We provide the first generalization error analysis for black-box learning through derivative-free optimization. Under the assumption of a Lipschitz and smooth unknown loss, we consider the Zeroth-order Stochastic Search (ZoSS) algorithm, that updates a $d$-dimensional model by replacing stochastic g…

Cited by 18SourcePDFScholar
2022

Learning Distributionally Robust Models at Scale via Composite Optimization

ICLR 2022poster

To train machine learning models that are robust to distribution shifts in the data, distributionally robust optimization (DRO) has been proven very effective. However, the existing approaches to learning a distributionally robust model either require solving complex optimization problems such as se…

Cited by 6SourcePDFScholar
2021

Federated Learning with Compression: Unified Analysis and Sharp Guarantees

AISTATS 2021poster

In federated learning, communication cost is often a critical bottleneck to scale up distributed optimization algorithms to collaboratively learn a model from millions of devices with potentially unreliable or limited communication and heterogeneous data distributions. Two notable trends to deal wit…

2019

Local SGD with Periodic Averaging: Tighter Analysis and Adaptive Synchronization

NeurIPS 2019poster

Communication overhead is one of the key challenges that hinders the scalability of distributed optimization algorithms. In this paper, we study local distributed SGD, where data is partitioned among computation nodes, and the computation nodes perform local updates with periodically exchanging the…

2019

Trading Redundancy for Communication: Speeding up Distributed SGD for Non-convex Optimization

ICML 2019oral

Communication overhead is one of the key challenges that hinders the scalability of distributed optimization algorithms to train large neural networks. In recent years, there has been a great deal of research to alleviate communication cost by compressing the gradient vector or using local updates a…

Cited by 90SourcePDFScholar