← Search

Philip Schroeder

2 accepted papers

2025

ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks

NeurIPS 2025poster

Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their utility in embodied settings, which require reasoning over lon…

Cited by 0SourceScholar
2025

THREAD: Thinking Deeper with Recursive Spawning

NAACL 2025long

Large language models (LLMs) have shown impressive capabilities across diverse settings, but still struggle as the length and complexity of the context increases. To address this challenge, we propose Thinking Recursively and Dynamically (ThReaD). THREAD frames model generation as a thread of execut…