2026
Revealing Differences in Multi-Modal Embeddings via Constrained Kernel Analysis
ICML 2026poster
Multi-modal representation models such as CLIP, SigLIP, and their variants are widely used to represent data across multiple modalities in modern learning systems. While these models are commonly evaluated through downstream performance, the analysis of their structural differences in how multi-moda…