2026
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
ICLR 2026poster
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than bi…