2026
Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning Framework
CVPR 2026
High-quality pixel-level responses remain a major bottleneck for multimodal large language models (MLLMs) in regional perception. Existing approaches generally attach regression decoders to MLLM features, achieving strong grounding performance but compromising end-to-end design and increasing traini