2025
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
EMNLP 2025
Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. However, our finding reveals they often mistakenly perceive text inputs lacking visual evidence as being part of the image, leading to erroneous responses. In light of thi