2026
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
CVPR 2026
Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Prior efforts to mitigate this have been limited, often defining these artifacts narrowly, for example, as simple high-norm