Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
Long-context models are essential for many applications but face inefficiencies in loading large KV caches during decoding. Prior methods enforce fixed token budgets for sparse attention, assuming a set number of tokens can approximate full attention. However, these methods overlook variations in th…