2024
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
ECCV 2024poster
"Most advanced visual grounding methods rely on Transformers for visual-linguistic feature fusion. However, these Transformer-based approaches encounter a significant drawback: the computational costs escalate quadratically due to the self-attention mechanism in the Transformer Encoder, particularly…