2026
RAR: Reversing Visual Attention Re-Sinking for Unlocking Potential in Multimodal Large Language Models
ICLR 2026poster
Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet they frequently exhibit suboptimal output layers, where intermediate decoder layers outperform the final ones, signaling underutilized model capacity. In this work, we delve into the root causes a…