2026
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
ICLR 2026poster
Vision Transformers have emerged as powerful, scalable and versatile representation learners. To capture both global and local features, a learnable [CLS] class token is typically prepended to the input sequence of patch tokens. Despite their distinct nature, both token types are processed identical…