AAAI 2026technical0 citations

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

Hui Huang, Yancheng He, Wei Liu, Muyun Yang, Jiaheng Liu, Kehai Chen, Bing Xu, Conghui Zhu

Abstract

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error

BibTeX
@inproceedings{aaai2026_longformrewardbe,
  title = {Long-form RewardBench: Evaluating Reward Models for Long-form Generation},
  author = {Hui Huang and Yancheng He and Wei Liu and Muyun Yang and Jiaheng Liu and Kehai Chen and Bing Xu and Conghui Zhu and Hailong Cao and Tiejun Zhao},
  booktitle = {AAAI 2026},
  year = {2026}
}