2026
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
ICML 2026spotlight
Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or step-level supervision, neither of which scales easily. In this work, we introduce a scalable method to …