← Search

Sujung Hong

1 accepted papers

2026

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models

ICML 2026poster

Large diffusion vision–language models (LDVLMs) have recently demonstrated competitive performance on multimodal tasks, emerging as a promising alternative to autoregressive models. They enable parallel decoding for efficient inference and leverage bidirectional attention to capture global context. …

Cited by 0SourceScholar