Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion Model
Qingyu Shi, Jinbin Bai, Zhuoran Zhao, Wenhao Chai, Kaidong Yu, Jianzong Wu, Shuangyong Song, Yunhai Tong
Abstract
Autoregressive unified models suffer from slow inference due to sequential decoding, and non-autoregressive unified models suffer from weak generalization due to limited pretrained backbones. We introduce Muddit, a unified discrete diffusion transformer that enables fast and parallel generation across both text and image modalities. Leveraging efficient token-level discrete denoising, strong visual priors, and a lightweight text decoder, Muddit supports flexible, high-quality generation with a compact architecture. Empirical results show that Muddit achieves competitive or superior performance compared to significantly larger AR-based models, in both quality and speed. The work highlights the potential of pure discrete diffusion as a scalable and effective backbone for multimodal generation. Code and models will be available.
BibTeX
@inproceedings{
shi2026beyond,
title={Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion Model},
author={Qingyu Shi and Jinbin Bai and Zhuoran Zhao and Wenhao Chai and Kaidong Yu and Jianzong Wu and Shuangyong Song and Yunhai Tong and Xiangtai Li and Xuelong Li and Shuicheng YAN},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=pG0WTde3pR}
}