2026
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Models
CVPR 2026
Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introduce DiG (Differential Grounding), a novel proxy task framework where MLLMs learn