Shaping Rewards for Reinforcement Learning with Imperfect Demonstrations using Generative Models
The potential benefits of model-free reinforcement learning to real robotics systems are limited by its uninformed exploration that leads to slow convergence, lack of data-efficiency, and unnecessary interactions with the environment. To address these drawbacks we propose a method that combines rein…