2025
SALS: Sparse Attention in Latent Space for KV Cache Compression
NeurIPS 2025poster
Large Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics withi…