← Search

Brahma S Pavse

5 accepted papers

2025

Stable Offline Value Function Learning with Bisimulation-based Representations

ICML 2025poster

In reinforcement learning, offline value function learning is the procedure of using an offline dataset to estimate the expected discounted return from each state when taking actions according to a fixed target policy. The stability of this procedure, i.e., whether it converges to its fixed-point, c…

Cited by 0SourcePDFScholar
2024

Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces

ICML 2024poster

In many reinforcement learning (RL) applications, we want policies that reach desired states and then keep the controlled system within an acceptable region around the desired states over an indefinite period of time. This latter objective is called *stability* and is especially important when the s…

2023

Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State Abstraction

AAAI 2023technical

We consider the problem of off-policy evaluation (OPE) in reinforcement learning (RL), where the goal is to estimate the performance of an evaluation policy, pie, using a fixed dataset, D, collected by one or more policies that may be different from pie. Current OPE algorithms may produce poor OPE e…

Cited by 7SourcePDFScholar
2023

State-Action Similarity-Based Representations for Off-Policy Evaluation

NeurIPS 2023poster

In reinforcement learning, off-policy evaluation (OPE) is the problem of estimating the expected return of an evaluation policy given a fixed dataset that was collected by running one or more different policies. One of the more empirically successful algorithms for OPE has been the fitted q-evaluati…

2020

RIDM: Reinforced Inverse Dynamics Modeling for Learning from a Single Observed Demonstration

RA-L 2020

Augmenting reinforcement learning with imitation learning is often hailed as a method by which to improve upon learning from scratch. However, most existing methods for integrating these two techniques are subject to several strong assumptions-chief among them that information about demonstrator act

Cited by 36SourceScholar