← Search

Li xiangtian

1 accepted papers

2026

Native Reasoning Models: Training Language Models to Reason on Unverifiable Data

ICLR 2026poster

The dominant paradigm for training large reasoning models—combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR)—is fundamentally constrained by its reliance on high-quality, human-annotated reasoning data and external verifiers. This dependency incurs sign…

Cited by 0SourceScholar