← Search

Xianwei Chen

2 accepted papers

2026

ReLaX: Reasoning with Latent Exploration for Large Reasoning Models

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated remarkable potential in enhancing the reasoning capability of Large Reasoning Models (LRMs). However, RLVR often drives the policy toward over-determinism, resulting in ineffective exploration and premature policy conver

Cited by 0SourcecodeScholar