2024
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
ECCV 2024poster
"Open-vocabulary semantic segmentation requires models to effectively integrate visual representations with open-vocabulary semantic labels. While Contrastive Language-Image Pre-training (CLIP) models shine in recognizing visual concepts from text, they often struggle with segment coherence due to t…