2025
ClusterAttn: KV Cache Compression under Intrinsic Attention Clustering
ACL 2025long
Sparse attention can effectively alleviate the significant demands on memory when large language models (LLMs) process long contexts. Existing methods typically apply the same sparse pattern across different attention heads and inputs. However, this uniform approach fails to capture the inherent div…