2024
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
EMNLP 2024main
Diffusion-based text-to-image models have demonstrated impressive achievements in diversity and aesthetics but struggle to generate images with legible visual texts. Existing backbone models have limitations such as misspelling, failing to generate texts, and lack of support for Chinese texts, but t…