Interpretable and Globally Optimal Prediction for Textual Grounding using Image Concepts
Raymond Yeh, Jinjun Xiong, Wen-Mei Hwu, Minh Do, Alexander Schwing
Abstract
Textual grounding is an important but challenging task for human-computer inter- action, robotics and knowledge mining. Existing algorithms generally formulate the task as selection from a set of bounding box proposals obtained from deep net based systems. In this work, we demonstrate that we can cast the problem of textual grounding into a unified framework that permits efficient search over all possible bounding boxes. Hence, the method is able to consider significantly more proposals and doesn’t rely on a successful first stage hypothesizing bounding box proposals. Beyond, we demonstrate that the trained parameters of our model can be used as word-embeddings which capture spatial-image relationships and provide interpretability. Lastly, at the time of submission, our approach outperformed the current state-of-the-art methods on the Flickr 30k Entities and the ReferItGame dataset by 3.08% and 7.77% respectively.
BibTeX
@inproceedings{NIPS2017_52292e0c,
author = {Yeh, Raymond and Xiong, Jinjun and Hwu, Wen-Mei and Do, Minh and Schwing, Alexander},
booktitle = {Advances in Neural Information Processing Systems},
editor = {I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Interpretable and Globally Optimal Prediction for Textual Grounding using Image Concepts},
url = {https://proceedings.neurips.cc/paper_files/paper/2017/file/52292e0c763fd027c6eba6b8f494d2eb-Paper.pdf},
volume = {30},
year = {2017}
}