2025
HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing
ICLR 2025poster
The frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leadin…