← Search

Gabriel Herbert Sarch

6 accepted papers

2025

Grounded Reinforcement Learning for Visual Reasoning

NeurIPS 2025poster

While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual reasoning introduces added complexity by requiring models to direct visual attention, interpret perceptual inputs, and ground abstract reasoning in spat…

Cited by 0SourcecodeScholar
2025

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

ACL 2025finding

A person’s demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences…

Cited by 0SourcePDFScholar
2025

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

EMNLP 2025

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce Disjoint-3DQA , a generative QA benchmark that evaluates this ability

Cited by 0SourcePDFScholar
2024

VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought

NeurIPS 2024spotlight

Large-scale generative language and vision-language models (LLMs and VLMs) excel in few-shot in-context learning for decision making and instruction following. However, they require high-quality exemplar demonstrations to be included in their context window. In this work, we ask: Can LLMs and VLMs g…

Cited by 5SourcePDFScholar
2023

Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural Images

NeurIPS 2023poster

The alignment between deep neural network (DNN) features and cortical responses currently provides the most accurate quantitative explanation for higher visual areas. At the same time, these model features have been critiqued as uninterpretable explanations, trading one black box (the human brain) f…

Cited by 9SourcePDFScholar
2023

Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models

EMNLP 2023long findings

Pre-trained and frozen LLMs can effectively map simple scene re-arrangement instructions to programs over a robot's visuomotor functions through appropriate few-shot example prompting. To parse open-domain natural language and adapt to a user's idiosyncratic procedures, not known during prompt engin…

Cited by 0SourcecodeScholar