← Search

Junlin Mu

1 accepted papers

2025

SALS: Sparse Attention in Latent Space for KV Cache Compression

NeurIPS 2025poster

Large Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics withi…

Cited by 0SourceScholar