← Search

Litong Feng*

1 accepted papers

2024

ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference

ECCV 2024poster

"Despite the success of large-scale pretrained Vision-Language Models (VLMs) especially CLIP in various open-vocabulary tasks, their application to semantic segmentation remains challenging, producing noisy segmentation maps with mis-segmented regions. In this paper, we carefully re-investigate the…