← Search

Zhuofu Chen

1 accepted papers

2025

TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention

ICLR 2025poster

Large language models (LLMs) have driven significant advancements across diverse NLP tasks, with long-context models gaining prominence for handling extended inputs. However, the expanding key-value (KV) cache size required by Transformer architectures intensifies the memory constraints, particularl…