← Search

Marcos Vinicius Treviso

3 accepted papers

2026

Long-Context Generalization with Sparse Attention

ICLR 2026poster

Transformer-based architectures traditionally employ softmax to compute attention weights, which produces dense distributions over all tokens in a sequence. While effective in many settings, this density has been shown to be detrimental for tasks that demand precise focus on fixed-size patterns…

Cited by 0SourcecodeScholar
2022

Learning to Scaffold: Optimizing Model Explanations for Teaching

NeurIPS 2022accept

Modern machine learning models are opaque, and as a result there is a burgeoning academic subfield on methods that explain these models' behavior. However, what is the precise goal of providing such explanations, and how can we demonstrate that explanations achieve this goal? Some research argues t…