2022
Dual Capsule Attention Mask Network with Mutual Learning for Visual Question Answering
COLING 2022main
A Visual Question Answering (VQA) model processes images and questions simultaneously with rich semantic information. The attention mechanism can highlight fine-grained features with critical information, thus ensuring that feature extraction emphasizes the objects related to the questions. However,…