2019
MUREL: Multimodal Relational Reasoning for Visual Question Answering
CVPR 2019poster
Multimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the visual content relevant to the question, this simple mechanism is arguably insufficient to model complex reasoning features…