AAAI 2026technical0 citations

ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion

Lishuai Gao, Jun-Yan He, Yingsen Zeng, Yujie Zhong, Xiaopeng Sun, Jie Hu, Zan Gao, Xiaoming Wei

Abstract

Current text-to-image models face challenges in visual text rendering: text encoders like CLIP and T5 lack glyph-level understanding and often struggle to distinguish between the specific words to be rendered and their intended semantic meaning within prompts. In addition, inconsistencies between the base model and its plugins further compromise the quality of synthesized images. In this paper, we enhance the existing text-to-image method by addressing the following aspects: (1) Text-Glyph Alignmentin a Visual Question Answering (VQA) manner to enable glyph understanding for the text encoder. This involves establishing an explicit alignment between the representations of the glyphs and their detailed attribute descriptions, which boosts the model

BibTeX
@inproceedings{aaai2026_vitypehighfideli,
  title = {ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion},
  author = {Lishuai Gao and Jun-Yan He and Yingsen Zeng and Yujie Zhong and Xiaopeng Sun and Jie Hu and Zan Gao and Xiaoming Wei},
  booktitle = {AAAI 2026},
  year = {2026}
}
ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion · AAAI 2026