2024
Dissecting Query-Key Interaction in Vision Transformers
NeurIPS 2024spotlight
Self-attention in vision transformers is often thought to perform perceptual grouping where tokens attend to other tokens with similar embeddings, which could correspond to semantically similar features of an object. However, attending to dissimilar tokens can be beneficial by providing contextual i…