2022
Convolutional Embedding Makes Hierarchical Vision Transformer Stronger
ECCV 2022poster
"Vision Transformers (ViTs) have recently dominated a range of computer vision tasks, yet it suffers from low training data efficiency and inferior local semantic representation capability without appropriate inductive bias. Convolutional neural networks (CNNs) inherently capture regional-aware sema…