← Search

Hedi Ben-younes

2 accepted papers

2019

MUREL: Multimodal Relational Reasoning for Visual Question Answering

CVPR 2019poster

Multimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the visual content relevant to the question, this simple mechanism is arguably insufficient to model complex reasoning features…

Cited by 385PDFcodeScholar
2017

MUTAN: Multimodal Tucker Fusion for Visual Question Answering

ICCV 2017poster

Bilinear models provide an appealing framework for mixing and merging information in Visual Question Answering (VQA) tasks. They help to learn high level associations between question meaning and visual concepts in the image, but they suffer from huge dimensionality issues. We introduce MUTAN, a mul…

Cited by 824PDFcodeScholar