NAACL 2022long72 citations
All You May Need for VQA are Image Captions
Soravit Changpinyo, Doron Kukliansy, Idan Szpektor, Xi Chen, Nan Ding, Radu Soricut
Abstract
Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we propose a method that automatically derives VQA examples at volume, by leveraging the abundance of existing image-caption annotations combined with neural models for textual question generation. We show that the resulting data is of high-quality. VQA models trained on our data improve state-of-the-art zero-shot accuracy by double digits and achieve a level of robustness that lacks in the same model trained on human-annotated VQA data.
BibTeX
@inproceedings{changpinyo-etal-2022-may,
title = "All You May Need for {VQA} are Image Captions",
author = "Changpinyo, Soravit and
Kukliansy, Doron and
Szpektor, Idan and
Chen, Xi and
Ding, Nan and
Soricut, Radu",
editor = "Carpuat, Marine and
de Marneffe, Marie-Catherine and
Meza Ruiz, Ivan Vladimir",
booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.naacl-main.142/",
doi = "10.18653/v1/2022.naacl-main.142",
pages = "1947--1963"
}