← Search

Shivansh Patel

7 accepted papers

2026

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

ICLR 2026poster

This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks—such as pouring, wiping, and mixing—purely by imitating AI-generated videos, without requiring any physical demonstrations or robot-specific training. Given a language…

Cited by 0SourcecodeScholar
2025

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

ICRA 2025

Task specification for robotic manipulation in open-world environments is challenging, requiring flexible and adaptive objectives that align with human intentions and can evolve through iterative feedback. We introduce Iterative Keypoint Reward (IKER), a visually grounded, Python-based reward functi

Cited by 1SourcecodeScholar
2021

Interpretation of Emergent Communication in Heterogeneous Collaborative Embodied Agents

ICCV 2021poster

Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of emergent forms of communication, we first introduce the collaborat…

Cited by 39PDFScholar
2021

Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments

EMNLP 2021main

In the Vision-and-Language Navigation (VLN) task an embodied agent navigates a 3D environment, following natural language instructions. A challenge in this task is how to handle ‘off the path’ scenarios where an agent veers from a reference path. Prior work supervises the agent with actions based on…

2020

MultiON: Benchmarking Semantic Map Memory using Multi-Object Navigation

NeurIPS 2020poster

Navigation tasks in photorealistic 3D environments are challenging because they require perception and effective planning under partial observability. Recent work shows that map-like memory is useful for long-horizon navigation tasks. However, a focused investigation of the impact of maps on navigat…

Cited by 120SourcePDFScholar
2019

U-CAM: Visual Explanation Using Uncertainty Based Class Activation Maps

ICCV 2019poster

Understanding and explaining deep learning models is an imperative task. Towards this, we propose a method that obtains gradient-based certainty estimates that also provide visual attention maps. Particularly, we solve for visual question answering task. We incorporate modern probabilistic deep lear…

Cited by 110PDFScholar