2025
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
ICCV 2025poster
Pre-trained vision-language models (VLMs), such as CLIP, have demonstrated impressive zero-shot recognition capability, but still underperform in dense prediction tasks. Self-distillation recently is emerging as a promising approach for fine-tuning VLMs to better adapt to local regions without requi…