2026
Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification
IJCAI 2026
Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-l