CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
"The popular CLIP model displays impressive zero-shot capabilities thanks to its seamless interaction with arbitrary text prompts. However, its lack of spatial awareness makes it unsuitable for dense computer vision tasks, e.g., semantic segmentation, without an additional fine-tuning step that ofte…