2026
QuRL: Rubrics As Judge For Open-Ended Question Answering
ICLR 2026poster
Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the performance of large language models (LLMs) on tasks with gold ground truth, such as code generation and mathematical reasoning. However, its application to open-ended question answering (QA) remains challenging, pr…