← Search

Kshiteej Sheth

2 accepted papers

2025

Improved Algorithms for Kernel Matrix-Vector Multiplication Under Sparsity Assumptions

ICLR 2025poster

Motivated by the problem of fast processing of attention matrices, we study fast algorithms for computing matrix-vector products for asymmetric Gaussian Kernel matrices $K\in \mathbb{R}^{n\times n}$. $K$'s columns are indexed by a set of $n$ keys $k_1,k_2\ldots, k_n\in \mathbb{R}^d$, rows by a set…

Cited by 0SourcePDFScholar
2025

Streaming Attention Approximation via Discrepancy Theory

NeurIPS 2025spotlight

Large language models (LLMs) have achieved impressive success, but their high memory requirements present challenges for long-context token generation. In this paper we study the streaming complexity of attention approximation, a key computational primitive underlying token generation. Our main…

Cited by 0SourceScholar