2026
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
ICML 2026poster
Vision Language Models (VLMs) are designed to extend Large Language Models (LLMs) with visual capabilities, yet in this work we observe a surprising phenomenon: VLMs can outperform their underlying LLMs on purely text-only tasks, particularly in long-context information retrieval. To investigate thi…