← Search

Sibo Li

1 accepted papers

2026

RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning

ICLR 2026poster

Despite rapid advancements in large language models (LLMs), the token-level autoregressive nature constrains their complex reasoning capabilities. To enhance LLM reasoning, inference-time techniques, including Chain/Tree/Graph-of-Thought(s), successfully improve the performance, as they are fairly c…

Cited by 0SourcecodeScholar