← Search

Uri Sherman

8 accepted papers

2025

Convergence of Policy Mirror Descent Beyond Compatible Function Approximation

ICML 2025poster

Modern policy optimization methods roughly follow the policy mirror descent (PMD) algorithmic template, for which there are by now numerous theoretical convergence results. However, most of these either target tabular environments, or can be applied effectively only when the class of policies be…

Cited by 0SourcePDFScholar
2025

Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime

NeurIPS 2025poster

We study population convergence guarantees of stochastic gradient descent (SGD) for smooth convex objectives in the interpolation regime, where the noise at optimum is zero or near zero. The behavior of the last iterate of SGD in this setting---particularly with large (constant) stepsizes---has rece…

Cited by 0SourceScholar
2025

Optimal Rates in Continual Linear Regression via Increasing Regularization

NeurIPS 2025poster

We study realizable continual linear regression under random task orderings, a common setting for developing continual learning theory. In this setup, the worst-case expected loss after $k$ learning iterations admits a lower bound of $\Omega(1/k)$. However, prior work using an unregularized scheme…

Cited by 0SourceScholar
2024

Rate-Optimal Policy Optimization for Linear Markov Decision Processes

ICML 2024oral

We study regret minimization in online episodic linear Markov Decision Processes, and propose a policy optimization algorithm that is computationally efficient, and obtains rate optimal $\widetilde O (\sqrt K)$ regret where $K$ denotes the number of episodes. Our work is the first to establish the o…

Cited by 12SourcePDFScholar
2023

Improved Regret for Efficient Online Reinforcement Learning with Linear Function Approximation

ICML 2023poster

We study reinforcement learning with linear function approximation and adversarially changing cost functions, a setup that has mostly been considered under simplifying assumptions such as full information feedback or exploratory conditions. We present a computationally efficient policy optimization…

Cited by 22SourcePDFScholar
2023

Regret Minimization and Convergence to Equilibria in General-sum Markov Games

ICML 2023poster

An abundance of recent impossibility results establish that regret minimization in Markov games with adversarial opponents is both statistically and computationally intractable. Nevertheless, none of these results preclude the possibility of regret minimization under the assumption that all parties…

Cited by 30SourcePDFScholar