2022
Improving Transformer with an Admixture of Attention Heads
NeurIPS 2022accept
Transformers with multi-head self-attention have achieved remarkable success in sequence modeling and beyond. However, they suffer from high computational and memory complexities for computing the attention matrix at each head. Recently, it has been shown that those attention matrices lie on a low-d…