2026
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
ICML 2026poster
As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2…