2024
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
ECCV 2024poster
"Highlighting particularly relevant regions of an image can improve the performance of vision-language models (VLMs) on various vision-language (VL) tasks by guiding the model to attend more closely to these regions of interest. For example, VLMs can be given a “visual prompt”, where visual markers…