DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge lies in inductive biases induced by distinct supervision signals: generation branch prefers high-fidelity, fine-grained represen…