2026
ActiveScope: Actively Seeking and Correcting Perception for MLLMs
ICML 2026poster
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding, yet they still struggle with fine-grained perception in high-resolution images. While existing training-free methods typically rely on attention-based localization or coarse-to-fine s…