2026
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
ICLR 2026poster
Although reinforcement learning with verifiable rewards (RLVR) shows promise in improving the reasoning ability of large language models (LLMs), the scaling up dilemma remains due to the reliance on human-annotated labels especially for complex tasks. Recent self-rewarding methods provide a label-fr…