SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformers
Over the past few years, vision transformers (ViTs) have consistently demonstrated remarkable performance across various visual recognition tasks. However, attempts to enhance their robustness have yielded limited success, mainly focusing on different training strategies, input patch augmentation, o…