← Search

Jiqian Dong

1 accepted papers

2026

Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement Learning

ICML 2026spotlight

Reinforcement Learning (RL) refines large language models (LLMs) by directly optimizing model behavior with reward signals. Although accurate state value estimation is essential for stable training in classical RL settings, it remains an understudied challenge in LLM post-training. In this work, we …

Cited by 0SourceScholar