Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursion
Vision Transformers (ViTs) have achieved remarkable progress in visual and multimodal tasks, yet their deployment remains costly. Token-adaptive methods reduce FLOPs through dynamic depth computation, but they face two limitations: (1) Global attention overemphasizes highly similar foreground region