2025
Streaming Attention Approximation via Discrepancy Theory
NeurIPS 2025spotlight
Large language models (LLMs) have achieved impressive success, but their high memory requirements present challenges for long-context token generation. In this paper we study the streaming complexity of attention approximation, a key computational primitive underlying token generation. Our main…