← Search

Dhruv Shah

41 accepted papers

2026

A Taxonomy for Evaluating Generalist Robot Manipulation Policies

ICRA 2026poster

Machine learning for robot manipulation promises to unlock generalization to novel tasks and environments. But how should we measure the progress of these policies towards generalization? Evaluating and quantifying generalization is the Wild West of modern robotics, with each work proposing and meas…

2026

LAP: Language-Action Pre-training Enables Zero-Shot Cross-Embodiment Transfer

RSS 2026poster

A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision–Language–Action models (VLAs) remain tightly coupled to their training embodiments and…

Cited by 0SourceScholar
2026

Learning to Drive Anywhere With Model-Based Reannotation

RA-L 2026

Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization.

Cited by 12SourcecodeScholar
2026

Learning to Drive Anywhere with Model-Based Reannotation

ICRA 2026poster

Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization. …

2026

Long-Context Robot Imitation Learning by Focusing on Key History Frames

RSS 2026poster

Many useful robot tasks require attending to the history of past observations. For example, finding an item in a room requires remembering which places have already been searched. However, the best-performing robot policies typically condition only on the current observation, limiting their applicab…

Cited by 0SourceScholar
2026

OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation

ICRA 2026poster

Humans can flexibly interpret and compose different goal specifications, such as language instructions, spatial coordinates, or visual references, when navigating to a destination. In contrast, most existing robotic navigation policies are trained on a single modality, limiting their adaptability to…

2026

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

RSS 2026poster

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is …

Cited by 26SourceScholar
2026

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

RSS 2026poster

Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision–Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VL…

Cited by 0SourceScholar
2026

Traversability-Aware Legged Navigation by Learning from Real-World Visual Data

ICRA 2026poster

The enhanced mobility brought by legged locomotion empowers quadrupedal robots to navigate through complex and unstructured environments. However, optimizing agile locomotion while accounting for the varying energy costs of traversing different terrains remains an open challenge. Most previous work …

2025

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

RSS 2025poster

In this work, we investigate how spatially-grounded auxiliary representations can provide both broad, high-level grounding, as well as direct, actionable information, and help policy learning performance and generalization. We study these mid-level representations across three critical dimensions: o…

Cited by 0PDFScholar
2025

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

CoRL 2025poster

How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in terms of predicting motion information from web data through human video generation and conditioning a robot policy on the generated video. Instead of…

Cited by 0SourceScholar
2025

Robot Data Curation with Mutual Information Estimators

RSS 2025poster

The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, litt…

Cited by 3PDFScholar
2025

STEER: Flexible Robotic Manipulation via Dense Language Grounding

ICRA 2025

The complexity of the real world demands robotic systems that can intelligently adapt to unseen situations. We present STEER, a robot learning framework that bridges highlevel, commonsense reasoning with precise, flexible low-level control. Our approach translates complex situational awareness into

Cited by 9SourcecodeScholar
2025

Vision Language Models are In-Context Value Learners

ICLR 2025spotlight

Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress estimator, or temporal value function, across different tasks and domains requires both a large amount of diverse data and methods which can s…

Cited by 2SourcePDFScholar
2024

GOAT: GO to Any Thing

RSS 2024poster

In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system cap…

2024

Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

CoRL 2024poster

An elusive goal in navigation research is to build an intelligent agent that can understand multimodal instructions including natural language and image, and perform useful navigation. To achieve this, we study a widely useful category of navigation tasks we call Multimodal Instruction Navigation wi…

Cited by 20SourceScholar
2024

NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration

ICRA 2024poster

Robotic learning for navigation in unfamiliar environments needs to provide policies for both task-oriented navigation (i.e., reaching a goal that the robot has located), and task-agnostic exploration (i.e., searching for a goal in a novel setting). Typically, these roles are handled by separate mod…

Cited by 124SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation

RSS 2024poster

Recent years in robotics and imitation learning have shown remarkable progress in training large-scale foundation models by leveraging data across a multitude of embodiments. The success of such policies might lead us to wonder: just how diverse can the robots in the training set be while still faci…

2024

SELFI: Autonomous Self-Improvement with RL for Vision-Based Navigation around People

CoRL 2024poster

Autonomous self-improving robots that interact and improve with experience are key to the real-world deployment of robotic systems. In this paper, we propose an online learning method, SELFI, that leverages online robot experience to rapidly fine-tune pre-trained control policies efficiently. SELFI…

Cited by 2SourceScholar
2023

AI Model Factory: Scaling AI for Industry 4.0 Applications

AAAI 2023technical

This demo paper discusses a scalable platform for emerging Data-Driven AI Applications targeted toward predictive maintenance solutions. We propose a common AI software architecture stack for building diverse AI Applications such as Anomaly Detection, Failure Pattern Analysis, Asset Health Forecasti…

Cited by 7SourcePDFScholar
2023

ExAug: Robot-Conditioned Navigation Policies via Geometric Experience Augmentation

ICRA 2023poster

Machine learning techniques rely on large and diverse datasets for generalization. Computer vision, natural language processing, and other applications can often reuse public datasets to train many different models. However, due to differences in physical configurations, it is challenging to leverag…

Cited by 23SourceScholar
2023

FastRLAP: A System for Learning High-Speed Driving via Deep RL and Autonomous Practicing

CoRL 2023poster

We present a system that enables an autonomous small-scale RC car to drive aggressively from visual observations using reinforcement learning (RL). Our system, FastRLAP, trains autonomously in the real world, without human interventions, and without requiring any simulation or expert demonstrations.…

Cited by 27SourceScholar
2023

GNM: A General Navigation Model to Drive Any Robot

ICRA 2023poster

Learning provides a powerful tool for vision-based navigation, but the capabilities of learning-based policies are constrained by limited training data. If we could combine data from all available sources, including multiple kinds of robots, we could train more powerful navigation models. In this pa…

Cited by 120SourcecodeScholar
2023

Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents

NeurIPS 2023poster

Recent progress in large language models (LLMs) has demonstrated the ability to learn and leverage Internet-scale knowledge through pre-training with autoregressive models. Unfortunately, applying such models to settings with embodied agents, such as robots, is challenging due to their lack of exper…

Cited by 142SourcePDFScholar
2023

Navigation with Large Language Models: Semantic Guesswork as a Heuristic for Planning

CoRL 2023poster

Navigation in unfamiliar environments presents a major challenge for robots: while mapping and planning techniques can be used to build up a representation of the world, quickly discovering a path to a desired goal in unfamiliar settings with such methods often requires lengthy mapping and explorati…

Cited by 118SourceScholar
2023

ViNT: A Foundation Model for Visual Navigation

CoRL 2023oral

General-purpose pre-trained models (``foundation models'') have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that are significantly smaller than those required for learning from scratch. Such models are typically trained on large and…

Cited by 154SourceScholar
2022

Hybrid Imitative Planning with Geometric and Predictive Costs in Off-road Environments

ICRA 2022poster

Geometric methods for solving open-world off-road navigation tasks, by learning occupancy and metric maps, provide good generalization but can be brittle in outdoor environments that violate their assumptions (e.g., tall grass). Learning-based methods can directly learn collision-free behavior from…

Cited by 17SourceScholar
2022

LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action

CoRL 2022poster

Goal-conditioned policies for robotic navigation can be trained on large, unannotated datasets, providing for good generalization to real-world settings. However, particularly in vision-based settings where specifying goals requires an image, this makes for an unnatural interface. Language provides…

Cited by 519SourcecodeScholar
2022

Offline Reinforcement Learning for Visual Navigation

CoRL 2022oral

Reinforcement learning can enable robots to navigate to distant goals while optimizing user-specified reward functions, including preferences for following lanes, staying on paved paths, or avoiding freshly mowed grass. However, online learning from trial-and-error for real-world robots is logistica…

Cited by 21SourcecodeScholar
2022

Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon Reasoning

ICLR 2022poster

Reinforcement learning can train policies that effectively perform complex tasks. However for long-horizon tasks, the performance of these methods degrades with horizon, often necessitating reasoning over and chaining lower-level skills. Hierarchical reinforcement learning aims to enable this by pro…

Cited by 41SourcePDFScholar
2021

Rapid Exploration for Open-World Navigation with Latent Goal Models

CoRL 2021oral

We describe a robotic learning system for autonomous exploration and navigation in diverse, open-world environments. At the core of our method is a learned latent variable model of distances and actions, along with a non-parametric topological memory of images. We use an information bottleneck to re…

Cited by 81SourceScholar
2021

ViNG: Learning Open-World Navigation with Visual Goals

ICRA 2021poster

We propose a learning-based navigation system for reaching visually indicated goals and demonstrate this system on a real mobile robot platform. Learning provides an appealing alternative to conventional methods for robotic navigation: instead of reasoning about environments in terms of geometry and…

Cited by 114SourceScholar
2020

Aerial Manipulation Using Hybrid Force and Position NMPC Applied to Aerial Writing

RSS 2020poster

Aerial manipulation aims at combining the maneuverability of aerial vehicles with the manipulation capabilities of robotic arms. This, however, comes at the cost of the additional control complexity due to the coupling of the dynamics of the two systems. In this paper we present a Nonlinear Model Pr…

Cited by 70SourcePDFScholar
2020

The Ingredients of Real World Robotic Reinforcement Learning

ICLR 2020spotlight

The success of reinforcement learning in the real world has been limited to instrumented laboratory scenarios, often requiring arduous human supervision to enable continuous learning. In this work, we discuss the required elements of a robotic system that can continually and autonomously improve wit…

Cited by 220SourceScholar