COLING 2022main2 citations

Belief Revision Based Caption Re-ranker with Visual Semantic Information

Ahmed Sabir, Francesc Moreno-Noguer, Pranava Madhyastha, Lluís Padró

Abstract

In this work, we focus on improving the captions generated by image-caption generation systems. We propose a novel re-ranking approach that leverages visual-semantic measures to identify the ideal caption that maximally captures the visual information in the image. Our re-ranker utilizes the Belief Revision framework (Blok et al., 2003) to calibrate the original likelihood of the top-n captions by explicitly exploiting semantic relatedness between the depicted caption and the visual context. Our experiments demonstrate the utility of our approach, where we observe that our re-ranker can enhance the performance of a typical image-captioning system without necessity of any additional training or fine-tuning.

BibTeX
@inproceedings{sabir-etal-2022-belief,
    title = "Belief Revision Based Caption Re-ranker with Visual Semantic Information",
    author = "Sabir, Ahmed  and
      Moreno-Noguer, Francesc  and
      Madhyastha, Pranava  and
      Padr{\'o}, Llu{\'i}s",
    editor = "Calzolari, Nicoletta  and
      Huang, Chu-Ren  and
      Kim, Hansaem  and
      Pustejovsky, James  and
      Wanner, Leo  and
      Choi, Key-Sun  and
      Ryu, Pum-Mo  and
      Chen, Hsin-Hsi  and
      Donatelli, Lucia  and
      Ji, Heng  and
      Kurohashi, Sadao  and
      Paggio, Patrizia  and
      Xue, Nianwen  and
      Kim, Seokhwan  and
      Hahm, Younggyun  and
      He, Zhong  and
      Lee, Tony Kyungil  and
      Santus, Enrico  and
      Bond, Francis  and
      Na, Seung-Hoon",
    booktitle = "Proceedings of the 29th International Conference on Computational Linguistics",
    month = oct,
    year = "2022",
    address = "Gyeongju, Republic of Korea",
    publisher = "International Committee on Computational Linguistics",
    url = "https://aclanthology.org/2022.coling-1.487/",
    pages = "5488--5506"
}
Belief Revision Based Caption Re-ranker with Visual Semantic Information · COLING 2022