2024
Vision Transformer with 2D Explicit Position Encoding
ICASSP 2024accepted
Recently, the Vision Transformer (ViT) has achieved outstanding performance in various computer vision tasks. Positional encoding is an indispensable component of ViT for handling the inherent structural information of images. However, attaching position encodings manually is a time-consuming proces…