2024
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
ECCV 2024poster
"Despite the success of large-scale pretrained Vision-Language Models (VLMs) especially CLIP in various open-vocabulary tasks, their application to semantic segmentation remains challenging, producing noisy segmentation maps with mis-segmented regions. In this paper, we carefully re-investigate the…