2024
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
ECCV 2024poster
"Recent approaches have shown that large-scale vision-language models such as CLIP can improve semantic segmentation performance. These methods typically aim for pixel-level vision-language alignment, but often rely on low-resolution image features from CLIP, resulting in class ambiguities along bou…