2026
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
ICLR 2026poster
Vision-Language Models (VLMs) achieve strong results on multimodal tasks such as visual question answering, yet they can still fail even when the correct visual evidence is present. In this work, we systematically investigate whether these failures arise from not perceiving the evidence or from not…