← Search

Aisha Urooj

4 accepted papers

2023

Learning Situation Hyper-Graphs for Video Question Answering

CVPR 2023poster

Answering questions about complex situations in videos requires not only capturing of the presence of actors, objects, and their relations, but also the evolution of these relationships over time. A situation hyper-graph is a representation that describes situations as scene sub-graphs for video fra…

2022

Weakly Supervised Grounding for VQA in Vision-Language Transformers

ECCV 2022poster

"Transformers for visual-language representation learning have been getting a lot of interest and shown tremendous performance on visual question answering (VQA) and grounding. However, most systems that show good performance of those tasks still rely on pre-trained object detectors during training,…

2021

Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using Capsules

CVPR 2021poster

The problem of grounding VQA tasks has seen an increased attention in the research community recently, with most attempts usually focusing on solving this task by using pretrained object detectors. However, pre-trained object detectors require bounding box annotations for detecting relevant objects…

Cited by 46PDFcodeScholar