2023
Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA
AAAI 2023technical
In this paper, we propose a novel multi-modal framework for Scene Text Visual Question Answering (STVQA), which requires models to read scene text in images for question answering. Apart from text or visual objects, which could exist independently, scene text naturally links text and visual modaliti…