← Search

Hugo Pitorro

3 accepted papers

2026

AdaSplash-2: Faster Differentiable Sparse Attention

ICML 2026poster

Sparse attention has been proposed as a way to alleviate the quadratic cost of transformers, a central bottleneck in long-context training. A promising line of work is $\alpha$-entmax attention, a differentiable sparse alternative to softmax that enables input-dependent sparsity yet has lagged behin…

Cited by 0SourceScholar
2026

Long-Context Generalization with Sparse Attention

ICLR 2026poster

Transformer-based architectures traditionally employ softmax to compute attention weights, which produces dense distributions over all tokens in a sequence. While effective in many settings, this density has been shown to be detrimental for tasks that demand precise focus on fixed-size patterns…

Cited by 0SourcecodeScholar