Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial Training
The Vision Transformer (ViT) has surpassed Convolutional Neural Networks (CNNs) in performance, becoming the de facto architecture in modern computer vision. However, despite its superior representational capacity, research on the adversarial robustness of ViTs remains limited, with most studies sti