← Search

Renan A. Rojas-Gomez

1 accepted papers

2024

Making Vision Transformers Truly Shift-Equivariant

CVPR 2024poster

In the field of computer vision Vision Transformers (ViTs) have emerged as a prominent deep learning architecture. Despite being inspired by Convolutional Neural Networks (CNNs) ViTs are susceptible to small spatial shifts in the input data - they lack shift-equivariance. To address this shortcoming…

Cited by 9SourcePDFScholar