2026
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
ICLR 2026poster
Efficient processing of high-resolution images is crucial for real-world vision–language applications. However, existing Large Vision-Language Models (LVLMs) incur substantial computational overhead due to the large number of vision tokens. With the advent of "thinking with images" models, reasoning…