← Search

Tu Yi

1 accepted papers

2025

HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing

ICLR 2025poster

The frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leadin…