2024
Investigating Compositional Challenges in Vision-Language Models for Visual Grounding
CVPR 2024highlight
Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performance gains contributed by large vision and language pre-training we find that state-of…