← Search

Tanmay Gangwani

11 accepted papers

2025

Selective Uncertainty Propagation in Offline RL

AAAI 2025technical

We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algorithms. To learn this, it is sufficient to evaluate the treatment effect of deviating from the behavioral policy at step…

Cited by 1SourcePDFScholar
2024

Multi-objective Optimization via Wasserstein-Fisher-Rao Gradient Flow

AISTATS 2024poster

Multi-objective optimization (MOO) aims to optimize multiple, possibly conflicting objectives with widespread applications. We introduce a novel interacting particle method for MOO inspired by molecular dynamics simulations. Our approach combines overdamped Langevin and birth-death dynamics, incorpo…

2022

Imitation Learning from Observations under Transition Model Disparity

ICLR 2022poster

Learning to perform tasks by leveraging a dataset of expert observations, also known as imitation learning from observations (ILO), is an important paradigm for learning skills without access to the expert reward function or the expert actions. We consider ILO in the setting where the expert and the…

2020

Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity

CoRL 2020

Quality-Diversity (QD) is a concept from Neuroevolution with some intriguing applications to Reinforcement Learning. It facilitates learning a population of agents where each member is optimized to simultaneously accumulate high task-returns and exhibit behavioral diversity compared to other members

2020

Mutual Information Based Knowledge Transfer Under State-Action Dimension Mismatch

UAI 2020poster

Deep reinforcement learning (RL) algorithms have achieved great success on a wide variety of sequential decision-making tasks. However, many of these algorithms suffer from high sample complexity when learning from scratch using environmental rewards, due to issues such as credit-assignment and high…

Cited by 28SourcePDFScholar
2019

Learning Belief Representations for Imitation Learning in POMDPs

UAI 2019poster

We consider the problem of imitation learning from expert demonstrations in partially observable Markov decision processes (POMDPs). Belief representations, which characterize the distribution over the latent states in a POMDP, have been modeled using recurrent neural networks and probabilistic late…