← Search

Xue Xu

2 accepted papers

2024

Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training

EMNLP 2024main

Diffusion-based text-to-image models have demonstrated impressive achievements in diversity and aesthetics but struggle to generate images with legible visual texts. Existing backbone models have limitations such as misspelling, failing to generate texts, and lack of support for Chinese texts, but t…

2024

UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion

ACL 2024long

Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific entities or scenes. This paper presents UNIMO-G, a simple multimo…