2026
The Extra Tokens Matter: Disentangled Representation Learning with Vision Transformers
ICML 2026poster
Vision Transformers increasingly incorporate extra tokens beyond patch tokens—from class tokens for aggregation to register tokens for artifact mitigation. While effective for their intended purposes, these tokens typically lack semantic structure. We ask a more ambitious question: Can we design reg…