2026
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
ICML 2026oral
Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently sequential decoding incurs substantial latency, motivating parallelization of the generation process. However, existing parallel reasoning approaches suffer from …