2025
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
NeurIPS 2025poster
Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their utility in embodied settings, which require reasoning over lon…