← Search

Filippo Utro

1 accepted papers

2025

Quantum Doubly Stochastic Transformers

NeurIPS 2025spotlight

At the core of the Transformer, the softmax normalizes the attention matrix to be right stochastic. Previous research has shown that this often de-stabilizes training and that enforcing the attention matrix to be doubly stochastic (through Sinkhorn’s algorithm) consistently improves performance acro…

Cited by 0SourceScholar