← Search

Kyle Jiang

1 accepted papers

2026

Stochastic Sparse Attention for Memory-Bound Inference

ICML 2026poster

Autoregressive decoding becomes bandwidth-limited at long contexts, as generating each token requires reading all $n_k$ key and value vectors from KV cache. We present Stochastic Additive No-mulT Attention (SANTA), a method that sparsifies value-cache access by sampling $S \ll n_k$ indices from the …

Cited by 0SourceScholar