← Search

Amar Kanakamedala

1 accepted papers

2026

RACE Attention: A Strictly Linear-Time Attention for Long-Sequence Training

ICLR 2026poster

Softmax Attention has a quadratic time complexity in sequence length, which becomes prohibitive to run at long contexts, even with highly optimized GPU kernels. For example, FlashAttention-2/3 (exact, GPU-optimized implementations of Softmax Attention) cannot complete a single forward–backward pass…

Cited by 0SourcecodeScholar