← Search

Qunli Zhang

1 accepted papers

2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

CVPR 2026

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we pre

Cited by 0SourcecodeScholar