2024
Exploring Question Guidance and Answer Calibration for Visually Grounded Video Question Answering
EMNLP 2024finding
Video Question Answering (VideoQA) tasks require not only correct answers but also visual evidence. The “localize-then-answer” strategy, while enhancing accuracy and interpretability, faces challenges due to the lack of temporal localization labels in VideoQA datasets. Existing methods often train t…