2025
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
NeurIPS 2025poster
Recent advances in inference-time compute have significantly improved performance on complex tasks by generating long chains of thought (CoTs) using Large Reasoning Models (LRMs). However, this improved accuracy comes at the cost of high inference latency due to the length of generated reasoning seq…