2025
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
NeurIPS 2025spotlight
Although applications involving long-context inputs are crucial for the effective utilization of large language models (LLMs), they also result in increased computational costs and reduced performance. To address this challenge, we propose an efficient, training-free prompt compression method that r…