VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions
Complex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language m…