2025
Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
AAAI 2025technical
Large language models (LLMs) have triggered a new stream of research focusing on compressing the context length to reduce the computational cost while ensuring the retention of helpful information for LLMs to answer the given question. Token-based removal methods are one of the most prominent approa…