2026
VIRTUE: Visual-Interactive Text-Image Universal Embedder
ICLR 2026poster
Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embedding models with instruction-following capabilities. However, existing embedding models lack visual-interactive capabil…