EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache Compression
The prohibitive memory footprint of the Key-Value (KV) cache imposes a critical bottleneck for efficient long-context LLM serving. Current compression techniques typically rely on static or uniform budget allocation, overlooking the significant heterogeneity in information density across attention h…