2026
Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement Learning
ICML 2026spotlight
Reinforcement Learning (RL) refines large language models (LLMs) by directly optimizing model behavior with reward signals. Although accurate state value estimation is essential for stable training in classical RL settings, it remains an understudied challenge in LLM post-training. In this work, we …