SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
Retrieval-augmented generation (RAG) has strong potential for producing accurate and factual outputs by combining language models (LMs) with evidence retrieved from large text corpora. However, current pipelines are limited by static chunking and flat retrieval: documents are split into short, prede…