← Search

Bowen Duan

1 accepted papers

2025

Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model

ICCV 2025poster

Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded…