2026
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
ICLR 2026poster
Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models typically rely on two key components: a Vision Transformer (ViT) and a Large Language Model (LLM). ViT encodes visual con…