← Search

Michael E. Sander

2 accepted papers

2022

Sinkformers: Transformers with Doubly Stochastic Attention

AISTATS 2022poster

Attention based models such as Transformers involve pairwise interactions between data points, modeled with a learnable attention matrix. Importantly, this attention matrix is normalized with the SoftMax operator, which makes it row-wise stochastic. In this paper, we propose instead to use Sinkhorn’…