2026
Closing the Modality Gap Aligns Group-Wise Semantics
ICLR 2026poster
In multimodal learning, CLIP has been recognized as the \textit{de facto} method for learning a shared latent space across multiple modalities, placing similar representations close to each other and moving them away from dissimilar ones. Although CLIP-based losses effectively align modalities at th…