NeurIPS 2018poster289 citations
Overcoming Language Priors in Visual Question Answering with Adversarial Regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, Stefan Lee
Abstract
Modern Visual Question Answering (VQA) models have been shown to rely heavily on superficial correlations between question and answer words learned during training -- \eg overwhelmingly reporting the type of room as kitchen or the sport being played as tennis, irrespective of the image. Most alarmingly, this shortcoming is often not well reflected during evaluation because the same strong priors exist in test distributions; however, a VQA system that fails to ground questions in image content would likely perform poorly in real-world settings.
BibTeX
@inproceedings{NEURIPS2018_67d96d45,
author = {Ramakrishnan, Sainandan and Agrawal, Aishwarya and Lee, Stefan},
booktitle = {Advances in Neural Information Processing Systems},
editor = {S. Bengio and H. Wallach and H. Larochelle and K. Grauman and N. Cesa-Bianchi and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Overcoming Language Priors in Visual Question Answering with Adversarial Regularization},
url = {https://proceedings.neurips.cc/paper_files/paper/2018/file/67d96d458abdef21792e6d8e590244e7-Paper.pdf},
volume = {31},
year = {2018}
}