The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
In reinforcement learning, specifying reward functions that capture the intended task can be very challenging. Reward learning aims to address this issue by *learning* the reward function. However, a learned reward model may have a low error on the data distribution, and yet subsequently produce a p…