2024
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
ECCV 2024poster
"Recent advances in contrastive language-image pretraining (CLIP) have demonstrated strong capabilities in zero-shot classification by aligning visual and textual features at an image level. However, in dense prediction tasks, CLIP often struggles to localize visual features within an image and fail…