2026
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
CVPR 2026
Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works have explored various combinations of vision encoders and LLMs, there still lacks a principled understanding of what makes a vision encoder suitable fo