2025
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
ICCV 2025poster
Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still struggles with fine-grained, region-level understanding, hindering its effectiveness on these dense prediction tasks. We…