← Search

WenZhang Wei

1 accepted papers

2026

Variational Adapter for Cross-modal Similarity Representation

ICML 2026poster

The core of vision-language models lies in measuring cross-modal similarity within a unified representation space. However, most image-text matching or multi-class image classification datasets lack fine-grained cross-modal matching annotations, forcing the continuous similarity space into binary cl…

Cited by 0SourceScholar