2024
One Shot Inverse Reinforcement Learning for Stochastic Linear Bandits
UAI 2024poster
The paradigm of inverse reinforcement learning (IRL) is used to specify the reward function of an agent purely from its actions and is critical for value alignment and AI safety. While IRL is successful in practice, theoretical guarantees remain nascent. Motivated by the need for IRL in large action…