AdanCA: Neural Cellular Automata As Adaptors For More Robust Vision Transformer
Vision Transformers (ViTs) demonstrate remarkable performance in image classification through visual-token interaction learning, particularly when equipped with local information via region attention or convolutions. Although such architectures improve the feature aggregation from different granular…