2021
Self-Supervised Online Reward Shaping in Sparse-Reward Environments
IROS 2021poster
We introduce Self-supervised Online Reward Shaping (SORS), which aims to improve the sample efficiency of any RL algorithm in sparse-reward environments by automatically densifying rewards. The proposed framework alternates between classification-based reward inference and policy update steps—the or…