UAI 2024poster1 citations

One Shot Inverse Reinforcement Learning for Stochastic Linear Bandits

Etash Guha, Jim James, Krishna Acharya, Vidya Muthukumar, Ashwin Pananjady

Abstract

The paradigm of inverse reinforcement learning (IRL) is used to specify the reward function of an agent purely from its actions and is critical for value alignment and AI safety. While IRL is successful in practice, theoretical guarantees remain nascent. Motivated by the need for IRL in large action spaces with limited data, we consider as a first step the problem of learning from a single sequence of actions (i.e., a demonstration) of a stochastic linear bandit algorithm. When the demonstrator employs the Phased Elimination algorithm, we develop a simple inverse learning procedure that estimates the linear reward function consistently in the time horizon with just a

BibTeX
@InProceedings{pmlr-v244-guha24a,
  title = 	 {One Shot Inverse Reinforcement Learning for Stochastic Linear Bandits},
  author =       {Guha, Etash and James, Jim and Acharya, Krishna and Muthukumar, Vidya and Pananjady, Ashwin},
  booktitle = 	 {Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence},
  pages = 	 {1491--1512},
  year = 	 {2024},
  editor = 	 {Kiyavash, Negar and Mooij, Joris M.},
  volume = 	 {244},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {15--19 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://raw.githubusercontent.com/mlresearch/v244/main/assets/guha24a/guha24a.pdf},
  url = 	 {https://proceedings.mlr.press/v244/guha24a.html},
  abstract = 	 {The paradigm of inverse reinforcement learning (IRL) is used to specify the reward function of an agent purely from its actions and is critical for value alignment and AI safety. While IRL is successful in practice, theoretical guarantees remain nascent. Motivated by the need for IRL in large action spaces with limited data, we consider as a first step the problem of learning from a single sequence of actions (i.e., a demonstration) of a stochastic linear bandit algorithm. When the demonstrator employs the Phased Elimination algorithm, we develop a simple inverse learning procedure that estimates the linear reward function consistently in the time horizon with just a
One Shot Inverse Reinforcement Learning for Stochastic Linear Bandits · UAI 2024