2026
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
ICLR 2026poster
Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek-R1) have led to remarkable improvements through long Chain-of-Thought (CoT). However, existing benchmarks mainly focus on immediate, single-horizon tasks, failing to adequately evaluate models’ ability to understand a…