2026
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
ICML 2026poster
Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-with-Images" methods alleviate this by iteratively zooming into regions of interes…