Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
From image-text pairs large-scale vision-language models (VLMs) learn to implicitly associate image regions with words which prove effective for tasks like visual question answering. However leveraging the learned association for open-vocabulary semantic segmentation remains a challenge. In this pap…