2026
Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoning
CVPR 2026
Multimodal large language models (MLLMs) have achieved remarkable success on diverse visual-language tasks. However, fixed-resolution models face challenges in perceiving fine-grained visual details, particularly due to *distracted attention* and *blurry vision*. To address these issues, we propose