Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks, but their inference remains computationally inefficient. We observe a common failure mode in many prevalent LLMs, overthinking, where models generate verbose and tangential reasoning traces even for simple quer…