2025
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
ICLR 2025poster
Large language models (LLMs) have driven significant advancements across diverse NLP tasks, with long-context models gaining prominence for handling extended inputs. However, the expanding key-value (KV) cache size required by Transformer architectures intensifies the memory constraints, particularl…