← Search

Ashwin Pananjady

6 accepted papers

2024

One Shot Inverse Reinforcement Learning for Stochastic Linear Bandits

UAI 2024poster

The paradigm of inverse reinforcement learning (IRL) is used to specify the reward function of an agent purely from its actions and is critical for value alignment and AI safety. While IRL is successful in practice, theoretical guarantees remain nascent. Motivated by the need for IRL in large action…

Cited by 1SourcePDFScholar
2023

Perceptual adjustment queries and an inverted measurement paradigm for low-rank metric learning

NeurIPS 2023poster

We introduce a new type of query mechanism for collecting human feedback, called the perceptual adjustment query (PAQ). Being both informative and cognitively lightweight, the PAQ adopts an inverted measurement scheme, and combines advantages from both cardinal and ordinal queries. We showcase the P…

2022

Learning from an Exploring Demonstrator: Optimal Reward Estimation for Bandits

AISTATS 2022poster

We introduce the “inverse bandit” problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator. Existing approaches to the related problem of inverse reinforcement learning assume the execution of an optimal policy, and thereby s…

2020

Preference learning along multiple criteria: A game-theoretic perspective

NeurIPS 2020poster

The literature on ranking from ordinal data is vast, and there are several ways to aggregate overall preferences from pairwise comparisons between objects. In particular, it is well-known that any Nash equilibrium of the zero-sum game induced by the preference matrix defines a natural solution conce…

Cited by 16SourcePDFScholar
2019

Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems

AISTATS 2019poster

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of a canonical stochastic, two-point, derivative-free method for linear-quadratic systems in which the initial state of the system is drawn at random. In partic…

Cited by 243SourcePDFScholar
2018

Gradient Diversity: a Key Ingredient for Scalable Distributed Learning

AISTATS 2018poster

It has been experimentally observed that distributed implementations of mini-batch stochastic gradient descent (SGD) algorithms exhibit speedup saturation and decaying generalization ability beyond a particular batch-size. In this work, we present an analysis hinting that high similarity between con…

Cited by 0SourcePDFScholar