← Search

Balasaravanan Thoravi Kumaravel

2 accepted papers

2025

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

ACL 2025finding

A person’s demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences…

Cited by 0SourcePDFScholar
2025

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

EMNLP 2025

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce Disjoint-3DQA , a generative QA benchmark that evaluates this ability

Cited by 0SourcePDFScholar