2020
Deep PQR: Solving Inverse Reinforcement Learning using Anchor Actions
ICML 2020accepted
We propose a reward function estimation framework for inverse reinforcement learning with deep energy-based policies. We name our method PQR, as it sequentially estimates the Policy, the Q-function, and the Reward function by deep learning. PQR does not assume that the reward solely depends on the s…