← Search

Ramki Gummadi

7 accepted papers

2024

Feasible $Q$-Learning for Average Reward Reinforcement Learning

AISTATS 2024poster

Average reward reinforcement learning (RL) provides a suitable framework for capturing the objective (i.e. long-run average reward) for continuing tasks, where there is often no natural way to identify a discount factor. However, existing average reward RL algorithms with sample complexity guarantee…

Cited by 6SourcePDFScholar
2024

Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation

ICML 2024spotlight

We prove that the combination of a target network and over-parameterized linear function approximation establishes a weaker convergence condition for bootstrapped value estimation in certain cases, even with off-policy data. Our condition is naturally satisfied for expected updates over the entire s…

2022

A Parametric Class of Approximate Gradient Updates for Policy Optimization

ICML 2022spotlight

Approaches to policy optimization have been motivated from diverse principles, based on how the parametric model is interpreted (e.g. value versus policy representation) or how the learning objective is formulated, yet they share a common goal of maximizing expected return. To better capture the com…

Cited by 0SourcePDFScholar
2022

Understanding and Leveraging Overparameterization in Recursive Value Estimation

ICLR 2022poster

The theory of function approximation in reinforcement learning (RL) typically considers low capacity representations that incur a tradeoff between approximation error, stability and generalization. Current deep architectures, however, operate in an overparameterized regime where approximation error…

Cited by 18SourcePDFScholar
2021

Characterizing the Gap Between Actor-Critic and Policy Gradient

ICML 2021spotlight

Actor-critic (AC) methods are ubiquitous in reinforcement learning. Although it is understood that AC methods are closely related to policy gradient (PG), their precise connection has not been fully characterized previously. In this paper, we explain the gap between AC and PG methods by identifying…

Cited by 21SourcePDFScholar
2019

Surrogate Objectives for Batch Policy Optimization in One-step Decision Making

NeurIPS 2019poster

We investigate batch policy optimization for cost-sensitive classification and contextual bandits---two related tasks that obviate exploration but require generalizing from observed rewards to action selections in unseen contexts. When rewards are fully observed, we show that the expected reward ob…

Cited by 34SourcePDFScholar
2018

Variational Rejection Sampling

AISTATS 2018poster

Learning latent variable models with stochastic variational inference is challenging when the approximate posterior is far from the true posterior, due to high variance in the gradient estimates. We propose a novel rejection sampling step that discards samples from the variational posterior which ar…

Cited by 0SourcePDFScholar