2026
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
ICML 2026poster
Large diffusion vision–language models (LDVLMs) have recently demonstrated competitive performance on multimodal tasks, emerging as a promising alternative to autoregressive models. They enable parallel decoding for efficient inference and leverage bidirectional attention to capture global context. …