EMNLP 2024main5 citations

Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation

Yuhui Zhang, Brandon McKinzie, Zhe Gan, Vaishaal Shankar, Alexander T Toshev

Abstract

Recent advances in image tokenizers, such as VQ-VAE, have enabled text-to-image generation using auto-regressive methods, similar to language modeling. However, these methods have yet to leverage pre-trained language models, despite their adaptability to various downstream tasks. In this work, we explore this gap by adapting a pre-trained language model for auto-regressive text-to-image generation, and find that pre-trained language models offer limited help. We provide a two-fold explanation by analyzing tokens from each modality. First, we demonstrate that image tokens possess significantly different semantics compared to text tokens, rendering pre-trained language models no more effective in modeling them than randomly initialized ones. Second, the text tokens in the image-text datasets are too simple compared to normal language model pre-training data, which causes the catastrophic degradation of language models’ capability.

BibTeX
@inproceedings{zhang-etal-2024-pre,
    title = "Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation",
    author = "Zhang, Yuhui  and
      McKinzie, Brandon  and
      Gan, Zhe  and
      Shankar, Vaishaal  and
      Toshev, Alexander T",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.75/",
    doi = "10.18653/v1/2024.emnlp-main.75",
    pages = "1281--1287"
}
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation · EMNLP 2024