2024
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
ECCV 2024oral
"Pre-trained vision-language models, e.g. CLIP, have been increasingly used to address the challenging Open-Vocabulary Segmentation (OVS) task, benefiting from their well-aligned vision-text embedding space. Typical solutions involve either freezing CLIP during training to unilaterally maintain its…