← Search

Kunzhao Xu

1 accepted papers

2026

ReVeal: Self-Evolving Code Agents via Reliable Self-Verification

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. Howerer, existing methods rely solely on outcome rewards, without explicitly optimizing verification or leveraging reliable signals from realistic environments, leading to unreliab…

Cited by 0SourceScholar