Visual Relationship Recognition via Language and Position Guided Attention
Hao Zhou, Chuanping Hu, Chongyang Zhang, Shengyang Shen
Abstract
Visual relationship recognition, as a challenging task used to distinguish the interactions between object pairs, has received much attention recently. Considering the fact that most visual relationships are semantic concepts defined by human beings, there are many human knowledge, or priors, hidden in them, which haven't been fully exploited by existing methods. In this work, we propose a novel visual relationship recognition model using language and position guided attention: language and position information are exploited and vectored firstly, and then both of them are used to guide the generation of attention maps. With the guided attention, the hidden human knowledge can be made better use to enhance the selection of spatial and channel features. Experiments on VRD [2] and VGR [1] show that, with language and position guided attention module, our proposed model achieves state-of-the-art performance.
BibTeX
@inproceedings{icassp2019_visualrelationsh,
title = {Visual Relationship Recognition via Language and Position Guided Attention},
author = {Hao Zhou and Chuanping Hu and Chongyang Zhang and Shengyang Shen},
booktitle = {ICASSP 2019},
year = {2019}
}