2026
Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
ICLR 2026oral
Active reasoning requires large language models (LLMs) to interact with external sources and strategically gather information to solve problems. Central to this process is belief tracking: maintaining a coherent understanding of the problem state and the missing information toward the solution. Howe…