← Search

Sohee Kim

1 accepted papers

2025

Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens

EMNLP 2025

Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. However, our finding reveals they often mistakenly perceive text inputs lacking visual evidence as being part of the image, leading to erroneous responses. In light of thi

Cited by 0SourcePDFScholar