← Search

Sarath Pattathil

5 accepted papers

2023

Revisiting the Linear-Programming Framework for Offline RL with General Function Approximation

ICML 2023poster

Offline reinforcement learning (RL) aims to find an optimal policy for sequential decision-making using a pre-collected dataset, without further interaction with the environment. Recent theoretical progress has focused on developing sample-efficient offline RL algorithms with various relaxed assumpt…

Cited by 27SourcePDFScholar
2023

Symmetric (Optimistic) Natural Policy Gradient for Multi-Agent Learning with Parameter Convergence

AISTATS 2023poster

Multi-agent interactions are increasingly important in the context of reinforcement learning, and the theoretical foundations of policy gradient methods have attracted surging research interest. We investigate the global convergence of natural policy gradient (NPG) algorithms in multi-agent learning…

Cited by 15SourcePDFScholar
2022

What is a Good Metric to Study Generalization of Minimax Learners?

NeurIPS 2022accept

Minimax optimization has served as the backbone of many machine learning problems. Although the convergence behavior of optimization algorithms has been extensively studied in minimax settings, their generalization guarantees, i.e., how the model trained on empirical data performs on the unseen test…

Cited by 16SourcePDFScholar
2020

A Unified Analysis of Extra-gradient and Optimistic Gradient Methods for Saddle Point Problems: Proximal Point Approach

AISTATS 2020poster

In this paper we consider solving saddle point problems using two variants of Gradient Descent-Ascent algorithms, Extra-gradient (EG) and Optimistic Gradient Descent Ascent (OGDA) methods. We show that both of these algorithms admit a unified analysis as approximations of the classical proximal poin…

Cited by 405SourcePDFScholar
2020

Tight last-iterate convergence rates for no-regret learning in multi-player games

NeurIPS 2020poster

We study the question of obtaining last-iterate convergence rates for no-regret learning algorithms in multi-player games. We show that the optimistic gradient (OG) algorithm with a constant step-size, which is no-regret, achieves a last-iterate rate of O(1/√T) with respect to the gap function in sm…

Cited by 115SourcePDFScholar