2026
Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Creative Writing
ICML 2026poster
Reinforcement learning with verifiable rewards (RLVR) on foundation models has led to significant improvements in math and code generation. Extending these gains to open-ended domains remains challenging: ground-truth verification is unavailable, human annotation is expensive, and learnt reward mode…