← Search

Prakash Panangaden

9 accepted papers

2026

Learning from Pairwise Preferences in Long-Term Decision Problems

ICML 2026poster

Agents that can beat or tie any other under a model of pairwise preference have strong guarantees for both user satisfaction and overall social welfare. However, searching for these agents in long-term decision problems is not computationally tractable with current approaches, which require the size…

Cited by 0SourceScholar
2025

Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning

ICLR 2025poster

Extracting relevant information from a stream of high-dimensional observations is a central challenge for deep reinforcement learning agents. Actor-critic algorithms add further complexity to this challenge, as it is often unclear whether the same information will be relevant to both the actor and t…

2024

Conditions on Preference Relations that Guarantee the Existence of Optimal Policies

AISTATS 2024poster

Learning from Preferential Feedback (LfPF) plays an essential role in training Large Language Models, as well as certain types of interactive learning agents. However, a substantial gap exists between the theory and application of LfPF algorithms. Current results guaranteeing the existence of optima…

Cited by 3SourcePDFScholar
2022

Continuous MDP Homomorphisms and Homomorphic Policy Gradient

NeurIPS 2022accept

Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definition of MDP homomorphisms to encompass continuous actions in continuous state spa…

2021

MICo: Improved representations via sampling-based state similarity for Markov decision processes

NeurIPS 2021poster

We present a new behavioural distance over the state space of a Markov decision process, and demonstrate the use of this distance as an effective means of shaping the learnt representations of deep reinforcement learning agents. While existing notions of state similarity are typically difficult to l…

2020

A Distributional Analysis of Sampling-Based Reinforcement Learning Algorithms

AISTATS 2020poster

We present a distributional approach to theoretical analyses of reinforcement learning algorithms for constant step-sizes. We demonstrate its effectiveness by presenting simple and unified proofs of convergence for a variety of commonly-used methods. We show that value-based methods such as TD(?) an…

Cited by 18SourcePDFScholar
2020

Latent Variable Modelling with Hyperbolic Normalizing Flows

ICML 2020poster

The choice of approximate posterior distributions plays a central role in stochastic variational inference (SVI). One effective solution is the use of normalizing flows \cut{defined on Euclidean spaces} to construct flexible posterior distributions. However, one key limitation of existing normalizin…

2015

Basis refinement strategies for linear value function approximation in MDPs

NeurIPS 2015poster

We provide a theoretical framework for analyzing basis function construction for linear value function approximation in Markov Decision Processes (MDPs). We show that important existing methods, such as Krylov bases and Bellman-error-based methods are a special case of the general framework we devel…

Cited by 8SourcePDFScholar