2025
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
ICLR 2025poster
Large language models (LLMs) augmented with retrieval exhibit robust performance and extensive versatility by incorporating external contexts. However, the input length grows linearly in the number of retrieved documents, causing a dramatic increase in latency. In this paper, we propose a novel para…