2025
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
CVPR 2025poster
This paper tackles the problem of video question answering (VideoQA), a task that often requires multi-step reasoning and a profound understanding of spatial-temporal dynamics. While large video-language models perform well on benchmarks, they often lack explainability and spatial-temporal grounding…