2025
It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data
CVPR 2025poster
The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar. This suggests that as foundation models mature, it may become possible to match…