2025
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
NeurIPS 2025poster
Vision Transformers (ViTs) have become a universal backbone for both image recognition and image generation. Yet their Multi–Head Self–Attention (MHSA) layer still performs a quadratic query–key interaction for \emph{every} token pair, spending the bulk of computation on visually weak or redundant…