← Search

Amitabh Swain

1 accepted papers

2026

Finding Distributed Object-Centric Properties in Self-Supervised Transformers

CVPR 2026

Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in \texttt [CLS] token attention maps of the final layer. However, these maps often contain spurious activations resulting in poor localization of objects. This is because the \textt

Cited by 0SourceScholar