← Search

Vikas Kapur

1 accepted papers

2025

Offloaded Reasoning: Efficient Inference for Large Language Models via Modular Reasoning and Refinement

EMNLP 2025

Large language models (LLMs) demonstrate strong reasoning capabilities but are expensive to run at inference time, limiting their practical deployment. We propose Offloaded Reasoning (OR), a modular strategy where a lightweight model generates intermediate reasoning traces that are then used by a la

Cited by 0SourcePDFScholar