2022
Weakly Supervised Grounding for VQA in Vision-Language Transformers
ECCV 2022poster
"Transformers for visual-language representation learning have been getting a lot of interest and shown tremendous performance on visual question answering (VQA) and grounding. However, most systems that show good performance of those tasks still rely on pre-trained object detectors during training,…