← Search

Brian Lu

1 accepted papers

2026

Generalization of RLVR Using Causal Reasoning as a Testbed

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for post-training large language models (LLMs) on complex reasoning tasks. Yet, the conditions under which RLVR yields robust generalization remain poorly understood. This paper provides an empirical study of R…

Cited by 0SourceScholar