2024
RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning
NeurIPS 2024poster
Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At its core, RLHF can be viewed as a specialized instance of Preference-based Reinforcement Learning (PbRL), where the pref…