2025
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
EMNLP 2025
Recent studies show that deep vision-only and language-only models—trained on disjoint modalities—nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of _where_ in each network this convergence emerges, _what_ visual or linguistic cues