2022
Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
CVPR 2022poster
Vision Transformers (ViTs) have emerged with superior performance on computer vision tasks compared to convolutional neural network (CNN)-based models. However, ViTs are mainly designed for image classification that generate single-scale low-resolution representations, which makes dense prediction t…