2024
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
CVPR 2024poster
Vision-Language Models (VLMs) such as Flamingo and GPT-4V have shown immense potential by integrating large language models with vision systems. Nevertheless these models face challenges in the fundamental computer vision task of object localisation due to their training on multimodal data containin…