2026
VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference
ICML 2026poster
Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP-based paradigm and harnesses the recent spatially-…