2026
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
CVPR 2026
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we pre