2025
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression
NAACL 2025findings
Retrieval-augmented generation (RAG) can supplement large language models (LLMs) by integrating external knowledge. However, as the number of retrieved documents increases, the input length to LLMs grows linearly, causing a dramatic increase in latency and a degradation in long-context understanding…