2023
Mulan: A Multi-Level Alignment Model for Video Question Answering
EMNLP 2023long findings
Video Question Answering (VideoQA) aims to answer questions about the visual content of a video. Current methods mainly focus on improving joint representations of video and text. However, these methods pay little attention to the fine-grained semantic interaction between video and text. In this pap…