Rating-Based Reinforcement Learning
Devin White, Mingkang Wu, Ellen Novoseller, Vernon J. Lawhern, Nicholas Waytowich, Yongcan Cao
Abstract
This paper develops a novel rating-based reinforcement learning approach that uses human ratings to obtain human guidance in reinforcement learning. Different from the existing preference-based and ranking-based reinforcement learning paradigms, based on human relative preferences over sample pairs, the proposed rating-based reinforcement learning approach is based on human evaluation of individual trajectories without relative comparisons between sample pairs. The rating-based reinforcement learning approach builds on a new prediction model for human ratings and a novel multi-class loss function. We conduct several experimental studies based on synthetic ratings and real human ratings to evaluate the effectiveness and benefits of the new rating-based reinforcement learning approach.
BibTeX
@article{White_Wu_Novoseller_Lawhern_Waytowich_Cao_2024, title={Rating-Based Reinforcement Learning}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/28886}, DOI={10.1609/aaai.v38i9.28886}, abstractNote={This paper develops a novel rating-based reinforcement learning approach that uses human ratings to obtain human guidance in reinforcement learning. Different from the existing preference-based and ranking-based reinforcement learning paradigms, based on human relative preferences over sample pairs, the proposed rating-based reinforcement learning approach is based on human evaluation of individual trajectories without relative comparisons between sample pairs. The rating-based reinforcement learning approach builds on a new prediction model for human ratings and a novel multi-class loss function. We conduct several experimental studies based on synthetic ratings and real human ratings to evaluate the effectiveness and benefits of the new rating-based reinforcement learning approach.}, number={9}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={White, Devin and Wu, Mingkang and Novoseller, Ellen and Lawhern, Vernon J. and Waytowich, Nicholas and Cao, Yongcan}, year={2024}, month={Mar.}, pages={10207-10215} }