← Search

Michael Friedlander

6 accepted papers

2026

Convergence Rate of the Last Iterate of Stochastic Proximal Algorithms

ICML 2026poster

We analyze two classical algorithms for solving additively composite convex optimization problems where the objective is the sum of a smooth term and a nonsmooth regularizer: proximal stochastic gradient method for a single regularizer; and the randomized incremental proximal method, which uses the …

Cited by 0SourceScholar
2024

Fair and Efficient Contribution Valuation for Vertical Federated Learning

ICLR 2024poster

Federated learning is an emerging technology for training machine learning models across decentralized data sources without sharing data. Vertical federated learning, also known as feature-based federated learning, applies to scenarios where data sources have the same sample IDs but different featur…

Cited by 49SourcePDFScholar
2021

Fast convergence of stochastic subgradient method under interpolation

ICLR 2021poster

This paper studies the behaviour of the stochastic subgradient descent (SSGD) method applied to over-parameterized nonsmooth optimization problems that satisfy an interpolation condition. By leveraging the composite structure of the empirical risk minimization problems, we prove that SSGD converges,…

Cited by 6SourcePDFScholar
2020

Greed Meets Sparsity: Understanding and Improving Greedy Coordinate Descent for Sparse Optimization

AISTATS 2020poster

We consider greedy coordinate descent (GCD) for composite problems with sparsity inducing regularizers, including 1-norm regularization and non-negative constraints. Empirical evidence strongly suggests that GCD, when initialized with the zero vector, has an implicit screening ability that usually s…

Cited by 20SourcePDFScholar
2020

Online mirror descent and dual averaging: keeping pace in the dynamic case

ICML 2020poster

Online mirror descent (OMD) and dual averaging (DA)—two fundamental algorithms for online convex optimization—are known to have very similar (and sometimes identical) performance guarantees when used with a fixed learning rate. Under dynamic learning rates, however, OMD is provably inferior to DA an…

Cited by 38SourcePDFScholar
2015

Coordinate Descent Converges Faster with the Gauss-Southwell Rule Than Random Selection

ICML 2015poster

There has been significant recent work on the theory and application of randomized coordinate descent algorithms, beginning with the work of  Nesterov [SIAM J. Optim., 22(2), 2012], who showed that a random-coordinate selection rule achieves the same convergence rate as the Gauss-Southwell selection…

Cited by 285SourcePDFScholar