2026
Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings
ICLR 2026poster
Vision–language models (VLMs) align images and text with remarkable success, yet the geometry of their shared embedding space remains poorly understood. To probe this geometry, we begin from the Iso-Energy Assumption, which exploits cross-modal redundancy: a concept that is truly shared should exhi…