2026
Uncovering Grounding IDs: How External Cues Shape Multi-Modal Binding
Amirmohammad Izadi, Hosein Hasani, Fatemeh Askari, Mobin Bagherian, Sadegh Mohammadian, Mohammad Izadi +1
ICML 2026poster
Large vision–language models (LVLMs) perform well on multimodal tasks, but their ability to reason and precisely align visual and textual information still has room for improvement. In this study, we show that external visual cues, such as symbols or grid lines, help LVLMs form more accurate connect…