2022
Reward Uncertainty for Exploration in Preference-based Reinforcement Learning
ICLR 2022poster
Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL methods are able to learn a more flexible reward model based on human preferences by actively incorporating human feedback, i.e. teacher's preferences between two clip…