2025
Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning
NeurIPS 2025poster
Reinforcement learning (RL) heavily depends on well-designed reward functions, which are often biased and difficult to design for complex behaviors. Preference-based RL (PbRL) addresses this by learning reward models from human feedback, but its practicality is constrained by a critical dilemma: wh…