← Search

Zeyuan Zhao

1 accepted papers

2026

Learning to Reason in Structured In-context Environments with Reinforcement Learning

ICLR 2026poster

Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental exploration. As the intrinsic properties of the environment determine the abilities that LLMs can learn, the environment plays a important role in the R…

Cited by 0SourceScholar