2026
DAMO: A DATA-EFFICIENT MULTIMODAL ORCHESTRATOR FOR TEMPORAL REASONING WITH VIDEO LLMS
ICASSP 2026oral
Large Language Models (LLMs) have recently been extended to the video domain, enabling sophisticated video-language understanding. However, existing Video LLMs often exhibit limitations in fine-grained temporal reasoning, restricting their ability to precisely attribute responses to specific video m…