IJCAI 20260 citations

Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification

Jiepeng Cai, Zhen Xu, Tiesong Zhao, Hau-San Wong, Si Wu

Abstract

Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-level semantic representations to modulate visual features, while overlooking the inherent structure of textual prompts. In this work, we propose a Syntactic Structure-guided Visual Grounding framework, referred to as SSVG. Specifically, to inject syntactic priors into unstructured visual representations, we design a semantic structure-based feature refinement module to adaptively modulate subject-centric and contextual visual features. To perform cross-modal alignment, we further incorporate a visual semantic consistency verification module, which leverages a subject-aware contrastive learning strategy to constrain and verify the semantic correspondence between the visual prediction and subject-level textual representation, thereby enhancing the model's robustness against semantically similar distractors. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art methods, and detailed analyses verify the effectiveness of each component.

Computer Vision: Multimodal learningComputer Vision: Recognition (object detection, categorization)Computer Vision: Segmentation, grouping and shape analysis
BibTeX
@inproceedings{ijcai2026_syntacticstructu,
  title = {Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification},
  author = {Jiepeng Cai and Zhen Xu and Tiesong Zhao and Hau-San Wong and Si Wu},
  booktitle = {IJCAI 2026},
  year = {2026}
}