2022
A free lunch from ViT: adaptive attention multi-scale fusion Transformer for fine-grained visual recognition
ICASSP 2022accepted
Learning subtle representation about object parts plays a vital role in fine-grained visual recognition (FGVR) field. The vision transformer (ViT) achieves promising results on computer vision due to its attention mechanism. Nonetheless, with the fixed size of patches in ViT, the class token in deep…