AAAI 2026technical0 citations

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains

Xuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo, Wentao Zhang

Abstract

Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consistency of model-generated responses and reference answers, since these responses are often lengthy, diverse, and nuanced. Rule-based verifiers struggle with complexity, prompting the use of model-based verifiers. Existing research primarily focuses on building better verifiers, yet a systematic evaluation of different types of verifiers

BibTeX
@inproceedings{aaai2026_verifybenchasyst,
  title = {VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains},
  author = {Xuzhao Li and Xuchen Li and Shiyu Hu and Yongzhen Guo and Wentao Zhang},
  booktitle = {AAAI 2026},
  year = {2026}
}