2026
RewardEval: Advancing Reward Model Evaluation
ICLR 2026poster
Reward models are used throughout the post-training of language models to capture nuanced signals from preference data and provide a training target for optimization across instruction following, reasoning, safety, and more domains. The community has begun establishing best practices for evaluating…