ICLR 2026poster0 citations

Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion Model

Qingyu Shi, Jinbin Bai, Zhuoran Zhao, Wenhao Chai, Kaidong Yu, Jianzong Wu, Shuangyong Song, Yunhai Tong

Abstract

Autoregressive unified models suffer from slow inference due to sequential decoding, and non-autoregressive unified models suffer from weak generalization due to limited pretrained backbones. We introduce Muddit, a unified discrete diffusion transformer that enables fast and parallel generation across both text and image modalities. Leveraging efficient token-level discrete denoising, strong visual priors, and a lightweight text decoder, Muddit supports flexible, high-quality generation with a compact architecture. Empirical results show that Muddit achieves competitive or superior performance compared to significantly larger AR-based models, in both quality and speed. The work highlights the potential of pure discrete diffusion as a scalable and effective backbone for multimodal generation. Code and models will be available.

discrete diffusionunified model
BibTeX
@inproceedings{
shi2026beyond,
title={Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion Model},
author={Qingyu Shi and Jinbin Bai and Zhuoran Zhao and Wenhao Chai and Kaidong Yu and Jianzong Wu and Shuangyong Song and Yunhai Tong and Xiangtai Li and Xuelong Li and Shuicheng YAN},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=pG0WTde3pR}
}
Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion Model · ICLR 2026