2026
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
ICLR 2026poster
Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. Howerer, existing methods rely solely on outcome rewards, without explicitly optimizing verification or leveraging reliable signals from realistic environments, leading to unreliab…