← Search

Hao Yi

2 accepted papers

2026

Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR

ICLR 2026poster

Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require large query budgets, making annotation costly. We investigate whether fewer but more informative queries can yield simila…

Cited by 0SourcecodeScholar
2025

SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin

EMNLP 2025

Enhancing the numerical and logical reasoning capabilities of Large Language Models (LLMs) has become a prominent research focus. Existing approaches exhibit notable limitations: inference-phase techniques, such as Chain of Thought, depend on prompt engineering and pretrained knowledge; sentence-lev

Cited by 0SourcePDFScholar