← Search

Ayush Shrivastava

8 accepted papers

2026

Point Prompting: Counterfactual Tracking with Video Diffusion Models

ICLR 2026poster

Recent advances in video generation have produced powerful diffusion models capable of generating high-quality, temporally coherent videos. We ask whether space-time tracking capabilities emerge automatically within these generators, as a consequence of the close connection between synthesizing and…

Cited by 0SourceScholar
2023

EXIF As Language: Learning Cross-Modal Associations Between Images and Camera Metadata

CVPR 2023highlight

We learn a visual representation that captures information about the camera that recorded a given photo. To do this, we train a multimodal embedding between image patches and the EXIF metadata that cameras automatically insert into image files. Our model represents this metadata by simply converting…

Cited by 15SourcePDFScholar
2022

TEACh: Task-Driven Embodied Agents That Chat

AAAI 2022technical

Robots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to c…

2022

VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

ACL 2022findings

Interactive robots navigating photo-realistic environments need to be trained to effectively leverage and handle the dynamic nature of dialogue in addition to the challenges underlying vision-and-language navigation (VLN). In this paper, we present VISITRON, a multi-modal Transformer-based navigator…

2020

Improving Vision-and-Language Navigation with Image-Text Pairs from the Web

ECCV 2020poster

Following a navigation instruction such as 'Walk down the stairs and stop at the brown sofa' requires embodied AI agents to ground referenced scene elements referenced (e.g. 'stairs') to visual content in the environment (pixels corresponding to 'stairs'). We ask the following question -- can we lev…

2020

Sim-to-Real Transfer for Vision-and-Language Navigation

CoRL 2020

We study the challenging problem of releasing a robot in a previously unseen environment, and having it follow unconstrained natural language navigation instructions. Recent work on the task of Vision-and-Language Navigation (VLN) has achieved significant progress in simulation. To assess the implic

2019

Chasing Ghosts: Instruction Following as Bayesian State Tracking

NeurIPS 2019poster

A visually-grounded navigation instruction can be interpreted as a sequence of expected observations and actions an agent following the correct trajectory would encounter and perform. Based on this intuition, we formulate the problem of finding the goal location in Vision-and-Language Navigation (VL…