2026
Learning to Reason in Structured In-context Environments with Reinforcement Learning
ICLR 2026poster
Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental exploration. As the intrinsic properties of the environment determine the abilities that LLMs can learn, the environment plays a important role in the R…