2025
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing
COLING 2025main
Visual grounding (VG) is an important task in vision and language that involves understanding the mutual relationship between query terms and images. However, existing VG datasets typically use simple and intuitive textual descriptions, with limited attribute and spatial information between images a…