2026
Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
ICLR 2026poster
Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on well-crafted task queries and corresponding ground-truth answers to provide accurate rewards, which requires significant human effort and hinders the sca…