2024
Look, Remember and Reason: Grounded Reasoning in Videos with Language Models
ICLR 2024poster
Multi-modal language models (LM) have recently shown promising performance in high-level reasoning tasks on videos. However, existing methods still fall short in tasks like causal or compositional spatiotemporal reasoning over actions, in which model predictions need to be grounded in fine-grained…