2018
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
CVPR 2018poster
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechan…