2026
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
CVPR 2026
Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in multimodal tasks.Despite their impressive performance, MLLMs suffer from the modality imbalance issue, where visual information is often underutilized compared to textual representations in deeper layers, leading to