2024
Efficient Vision Transformers with Partial Attention
ECCV 2024poster
"As a core of Vision Transformer (ViT), self-attention has high versatility in modeling long-range spatial interactions because every query attends to all spatial locations. Although ViT achieves promising performance in visual tasks, self-attention’s complexity is quadratic with token lengths. This…