← Search

Ruijin Ding

2 accepted papers

2026

Search Self-Play: Pushing the Frontier of Agent Capability without Supervision

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on well-crafted task queries and corresponding ground-truth answers to provide accurate rewards, which requires significant human effort and hinders the sca…

Cited by 0SourcecodeScholar
2023

Adversarial Counterfactual Environment Model Learning

NeurIPS 2023spotlight

An accurate environment dynamics model is crucial for various downstream tasks in sequential decision-making, such as counterfactual prediction, off-policy evaluation, and offline reinforcement learning. Currently, these models were learned through empirical risk minimization (ERM) by step-wise fit…