2025
Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question Answering
ICML 2025poster
Multimodal Large Language Models (Multimodal LLMs) have shown their strength in Video Question Answering (VideoQA). However, due to the black-box nature of end-to-end training strategies, existing approaches based on Multimodal LLMs suffer from the lack of interpretability for VideoQA: they can neit…