2022
How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers
EMNLP 2022finding
The attention mechanism is considered the backbone of the widely-used Transformer architecture. It contextualizes the input by computing input-specific attention matrices. We find that this mechanism, while powerful and elegant, is not as important as typically thought for pretrained language models…