← Search

Yongzhen Guo

1 accepted papers

2026

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains

AAAI 2026technical

Large language models (LLMs) increasingly rely on reinforcement learning (RL) to enhance their reasoning capabilities through feedback. A critical challenge is verifying the consistency of model-generated responses and reference answers, since these responses are often lengthy, diverse, and nuanced.

Cited by 0SourcePDFScholar