← Search

Jack Cai

2 accepted papers

2026

h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning

ICML 2026spotlight

Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or step-level supervision, neither of which scales easily. In this work, we introduce a scalable method to …

Cited by 0SourceScholar