2023
SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation
ICML 2023poster
Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowled…