← Search

Siyue Li

1 accepted papers

2025

VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions

ACL 2025long

Complex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language m…