2026
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
ICLR 2026poster
Unified Vision–Language Models (UVLMs) aim to advance multimodal learning by supporting both understanding and generation within a single framework. However, existing approaches largely focus on architectural unification while overlooking the need for explicit interaction between the two capabilitie…