2025
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
NeurIPS 2025poster
Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand, and reason about dynamic real-world scenarios. However, existing Multimodal Large Language Models (MLLMs) struggle with simultaneously ensuring the a…