← Search

Gaurav Sukhatme

16 accepted papers

2026

Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation

ICRA 2026poster

Reinforcement learning has enabled significant progress in complex domains such as coordinating and navigating multiple quadrotors. However, even well-trained policies remain vulnerable to collisions in obstacle-rich environments. Addressing these infrequent but critical safety failures through retr…

2026

Learning Geometry-Aware Nonprehensile Pushing and Pulling with Dexterous Hands

ICRA 2026poster

Nonprehensile manipulation, such as pushing and pulling, enables robots to move, align, or reposition objects that may be difficult to grasp due to their geometry, size, or relationship to the robot or the environment. Much of the existing work in nonprehensile manipulation relies on parallel-jaw gr…

2026

ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation

ICRA 2026poster

Training robust bimanual manipulation policies via imitation learning requires demonstration data with broad coverage over robot poses, contacts, and scene contexts. However, collecting diverse and precise real-world demonstrations is costly and time-consuming, which hinders scalability. Prior works…

2026

Refinery: Active Fine-Tuning and Deployment-Time Optimization for Contact-Rich Policies

ICRA 2026poster

Simulation-based learning has enabled policies for precise, contact-rich tasks (e.g., robotic assembly) to reach high success rates (~80%) under high levels of observation noise and control error. Although such performance may be sufficient for research applications, it falls short of industry stand…

2024

VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation

AAAI 2024technical

Outdoor Vision-and-Language Navigation (VLN) requires an agent to navigate through realistic 3D outdoor environments based on natural language instructions. The performance of existing VLN methods is limited by insufficient diversity in navigation environments and limited training data. To address t…

Cited by 8SourcePDFScholar
2021

Sampling-Based Motion Planning on Sequenced Manifolds

RSS 2021poster

We address the problem of planning robot motions in constrained configuration spaces where the constraints change throughout the motion. The problem is formulated as a fixed sequence of intersecting manifolds; which the robot needs to traverse in order to solve the task. We specify a class of sequen…

2020

Learning Equality Constraints for Motion Planning on Manifolds

CoRL 2020

Constrained robot motion planning is a widely used technique to solve complex robot tasks. We consider the problem of learning representations of constraints from demonstrations with a deep neural network, which we call Equality Constraint Manifold Neural Network (ECoMaNN). The key idea is to learn

2020

Motion Planner Augmented Reinforcement Learning for Robot Manipulation in Obstructed Environments

CoRL 2020

Deep reinforcement learning (RL) agents are able to learn contact-rich manipulation tasks by maximizing a reward signal, but require large amounts of experience, especially in environments with many obstacles that complicate exploration. In contrast, motion planners use explicit models of the agent

Cited by 0SourcePDFScholar
2020

Never Stop Learning: The Effectiveness of Fine-Tuning in Robotic Reinforcement Learning

CoRL 2020

One of the great promises of robot learning systems is that they will be able to learn from their mistakes and continuously adapt to ever-changing environments. Despite this potential, most of the robot learning systems today produce static policies that are not further adapted during deployment, be

Cited by 0SourcePDFScholar
2020

Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning

ICML 2020poster

Increasing the scale of reinforcement learning experiments has allowed researchers to achieve unprecedented results in both training sophisticated agents for video games, and in sim-to-real transfer for robotics. Typically such experiments rely on large distributed systems and require expensive hard…

2019

Coordinating multi-robot systems through environment partitioning for adaptive informative sampling

ICRA 2019poster

As robotic platforms have become more capable and autonomous, they have increasingly been utilized in time sensitive applications such as search and rescue. To that end, we have developed a system for teams of robots to efficiently explore an environment while taking sensor measurements. The system…

Cited by 40SourceScholar
2018

Accelerating Goal-Directed Reinforcement Learning by Model Characterization

IROS 2018poster

We propose a hybrid approach aimed at improving the sample efficiency in goal-directed reinforcement learning. We do this via a two-step mechanism where firstly, we approximate a model from Model-Free reinforcement learning. Then, we leverage this approximate model along with a notion of reachabilit…

Cited by 3SourceScholar
2018

Solving Markov Decision Processes with Reachability Characterization from Mean First Passage Times

IROS 2018poster

A new mechanism for efficiently solving the Markov decision processes (MDPs) is proposed in this paper. We introduce the notion of reachability landscape where we use the Mean First Passage Time (MFPT) as a means to characterize the reachability of every state in the state space. We show that such r…

Cited by 6SourceScholar
2017

Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

ICML 2017poster

Reinforcement learning algorithms for real-world robotic applications must be able to handle complex, unknown dynamical systems while maintaining data-efficient learning. These requirements are handled well by model-free and model-based RL approaches, respectively. In this work, we aim to combine th…

Cited by 227SourcePDFScholar
2017

Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets

NeurIPS 2017poster

Imitation learning has traditionally been applied to learn a single task from demonstrations thereof. The requirement of structured and isolated demonstrations limits the scalability of imitation learning approaches as they are difficult to apply to real-world scenarios, where robots have to be able…

Cited by 202SourcePDFScholar
2017

Trajectory Optimization for Self-Calibration and Navigation

RSS 2017poster

Trajectory generation approaches for mobile robots generally aim to optimize with respect to a cost function such as energy, execution time, or other mission-relevant parameters within the constraints of vehicle dynamics and obstacles in the environment. We propose to add the cost of state observabi…

Cited by 35SourcePDFScholar