← Search

Marcos V. Treviso

3 accepted papers

2026

AdaSplash-2: Faster Differentiable Sparse Attention

ICML 2026poster

Sparse attention has been proposed as a way to alleviate the quadratic cost of transformers, a central bottleneck in long-context training. A promising line of work is $\alpha$-entmax attention, a differentiable sparse alternative to softmax that enables input-dependent sparsity yet has lagged behin…

Cited by 0SourceScholar
2024

xTower: A Multilingual LLM for Explaining and Correcting Translation Errors

EMNLP 2024finding

While machine translation (MT) systems are achieving increasingly strong performance on benchmarks, they often produce translations with errors and anomalies. Understanding these errors can potentially help improve the translation quality and user experience. This paper introduces xTower, an open la…

Cited by 5SourcePDFScholar