COLING 2020main10 citations

Image Caption Generation for News Articles

Zhishen Yang, Naoaki Okazaki

Abstract

In this paper, we address the task of news-image captioning, which generates a description of an image given the image and its article body as input. This task is more challenging than the conventional image captioning, because it requires a joint understanding of image and text. We present a Transformer model that integrates text and image modalities and attends to textual features from visual features in generating a caption. Experiments based on automatic evaluation metrics and human evaluation show that an article text provides primary information to reproduce news-image captions written by journalists. The results also demonstrate that the proposed model outperforms the state-of-the-art model. In addition, we also confirm that visual features contribute to improving the quality of news-image captions.

BibTeX
@inproceedings{yang-okazaki-2020-image,
    title = "Image Caption Generation for News Articles",
    author = "Yang, Zhishen  and
      Okazaki, Naoaki",
    editor = "Scott, Donia  and
      Bel, Nuria  and
      Zong, Chengqing",
    booktitle = "Proceedings of the 28th International Conference on Computational Linguistics",
    month = dec,
    year = "2020",
    address = "Barcelona, Spain (Online)",
    publisher = "International Committee on Computational Linguistics",
    url = "https://aclanthology.org/2020.coling-main.176/",
    doi = "10.18653/v1/2020.coling-main.176",
    pages = "1941--1951"
}
Image Caption Generation for News Articles · COLING 2020