2026
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
IJCAI 2026
While Vision-Language Models (VLMs) have shown promise in textual understanding, they face significant challenges when handling long context and complex reasoning tasks. In this paper, we dissect the internal mechanisms governing long-context processing in VLMs to understand their performance bottle