2025
Offloaded Reasoning: Efficient Inference for Large Language Models via Modular Reasoning and Refinement
EMNLP 2025
Large language models (LLMs) demonstrate strong reasoning capabilities but are expensive to run at inference time, limiting their practical deployment. We propose Offloaded Reasoning (OR), a modular strategy where a lightweight model generates intermediate reasoning traces that are then used by a la