EMNLP 2024main4 citations

Altogether: Image Captioning via Re-aligning Alt-text

Hu Xu, Po-Yao Huang, Xiaoqing Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy

Abstract

This paper focuses on creating synthetic data to improve the quality of image captions. Existing works typically have two shortcomings. First, they caption images from scratch, ignoring existing alt-text metadata, and second, lack transparency if the captioners’ training data (e.g. GPT) is unknown. In this paper, we study a principled approach Altogether based on the key idea to edit and re-align existing alt-texts associated with the images. To generate training data, we perform human annotation where annotators start with the existing alt-text and re-align it to the image content in multiple rounds, consequently constructing captions with rich visual concepts. This differs from prior work that carries out human annotation as a one-time description task solely based on images and annotator knowledge. We train a captioner on this data that generalizes the process of re-aligning alt-texts at scale. Our results show our Altogether approach leads to richer image captions that also improve text-to-image generation and zero-shot image classification tasks.

BibTeX
@inproceedings{xu-etal-2024-altogether,
    title = "Altogether: Image Captioning via Re-aligning Alt-text",
    author = "Xu, Hu  and
      Huang, Po-Yao  and
      Tan, Xiaoqing  and
      Yeh, Ching-Feng  and
      Kahn, Jacob  and
      Jou, Christine  and
      Ghosh, Gargi  and
      Levy, Omer  and
      Zettlemoyer, Luke  and
      Yih, Wen-tau  and
      Li, Shang-Wen  and
      Xie, Saining  and
      Feichtenhofer, Christoph",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1075/",
    doi = "10.18653/v1/2024.emnlp-main.1075",
    pages = "19302--19318"
}
Altogether: Image Captioning via Re-aligning Alt-text · EMNLP 2024