2026
Reward Model Evaluation via Automatically-Ranked Policy Alignment
AAAI 2026technical
Evaluating reward models is a fundamental challenge in Reinforcement Learning (RL), particularly in settings where the reward model is learned or manually designed. The standard paradigm for Reward Model Evaluation (RME) involves training an optimal policy via RL on the given reward model and assess