← Search

Jihang Zhang

2 accepted papers

2025

HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing

ICLR 2025poster

The frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leadin…

2025

SALS: Sparse Attention in Latent Space for KV Cache Compression

NeurIPS 2025poster

Large Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics withi…

Cited by 0SourceScholar