2025
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
ICCV 2025poster
Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded…