← Search

Zoe Wanying He

1 accepted papers

2025

Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models

EMNLP 2025

Recent studies show that deep vision-only and language-only models—trained on disjoint modalities—nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of _where_ in each network this convergence emerges, _what_ visual or linguistic cues

Cited by 0SourcePDFScholar