2025
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
ICCV 2025poster
Multi-head self-attention (MHSA) is a key component of Transformers, a widely popular architecture in both language and vision. Multiple heads intuitively enable different parallel processes over the same input. Yet, they also obscure the attribution of each input patch to the output of a model. We…