← Search

Ankur Handa

31 accepted papers

2026

DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands

ICRA 2026poster

One of the most important, yet challenging, skills for a dexterous robot is grasping a diverse range of objects. Much of the prior work has been limited by speed, generality, or reliance on depth maps and object poses. In this paper, we introduce DextrAH-RGB, a system that can perform dexterous arm-…

2025

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

CVPR 2025poster

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLA…

2025

FORGE: Force-Guided Exploration for Robust Contact-Rich Manipulation Under Uncertainty

RA-L 2025

We present FORGE, a method for sim-to-real transfer of force-aware manipulation policies in the presence of significant pose uncertainty. During simulation-based policy learning, FORGE combines a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">force

Cited by 31SourceScholar
2025

Synthetica: Large Scale Synthetic Data Generation for Robot Perception

IROS 2025

Vision-based object detectors are a crucial basis for robotics applications as they provide valuable information about object localization in the environment. These need to ensure high reliability in different lighting conditions, occlusions, and visual artifacts, all while running in real-time. Col

Cited by 6SourceScholar
2024

AutoMate: Specialist and Generalist Assembly Policies over Diverse Geometries

RSS 2024poster

Robotic assembly for high-mixture settings requires adaptivity to diverse parts and poses, which is an open challenge. Meanwhile, in other areas of robotics, large models and sim-to-real have led to tremendous progress. Inspired by such work, we present AutoMate, a learning framework and system that…

Cited by 15SourcePDFScholar
2024

DextrAH-G: Pixels-to-Action Dexterous Arm-Hand Grasping with Geometric Fabrics

CoRL 2024poster

A pivotal challenge in robotics is achieving fast, safe, and robust dexterous grasping across a diverse range of objects, an important goal within industrial applications. However, existing methods often have very limited speed, dexterity, and generality, along with limited or no hardware safety gua…

Cited by 13SourceScholar
2024

Geometric Fabrics: a Safe Guiding Medium for Policy Learning

ICRA 2024poster

Robotics policies are always subjected to complex, second order dynamics that entangle their actions with resulting states. In reinforcement learning (RL) contexts, policies have the burden of deciphering these complicated interactions over massive amounts of experience and complex reward functions…

Cited by 6SourceScholar
2023

CuRobo: Parallelized Collision-Free Robot Motion Generation

ICRA 2023poster

This paper explores the problem of collision-free motion generation for manipulators by formulating it as a global motion optimization problem. We develop a parallel optimization technique to solve this problem and demonstrate its effectiveness on massively parallel GPUs. We show that combining simp…

Cited by 74SourceScholar
2023

DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality

ICRA 2023poster

Recent work has demonstrated the ability of deep reinforcement learning (RL) algorithms to learn complex robotic behaviours in simulation, including in the domain of multi-fingered manipulation. However, such models can be challenging to transfer to the real world due to the gap between simulation a…

Cited by 146SourceScholar
2023

DexPBT: Scaling up Dexterous Manipulation for Hand-Arm Systems with Population Based Training

RSS 2023poster

In this work, we propose algorithms and methods that enable learning dexterous object manipulation using simulated one- or two-armed robots equipped with multi-fingered hand end-effectors. Using a parallel GPU-accelerated physics simulator (Isaac Gym), we implement challenging tasks for these robots…

2023

Imitating Task and Motion Planning with Visuomotor Transformers

CoRL 2023poster

Imitation learning is a powerful tool for training robot manipulation policies, allowing them to learn from expert demonstrations without manual programming or trial-and-error. However, common methods of data collection, such as human supervision, scale poorly, as they are time-consuming and labor-i…

Cited by 56SourcecodeScholar
2023

IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality

RSS 2023poster

Robotic assembly is a longstanding challenge, requiring contact-rich interaction and high precision and accuracy. Many applications also require adaptivity to diverse parts, poses, and environments, as well as low cycle times. In other areas of robotics, simulation is a powerful tool to develop algo…

2022

Factory: Fast Contact for Robotic Assembly

RSS 2022poster

Robotic assembly is one of the oldest and most challenging applications of robotics. In other areas of robotics, such as perception and grasping, simulation has rapidly accelerated research progress, particularly when combined with modern deep learning. However, accurately, efficiently, and robustly…

2022

Neural Geometric Fabrics: Efficiently Learning High-Dimensional Policies from Demonstration

CoRL 2022poster

Learning dexterous manipulation policies for multi-fingered robots has been a long-standing challenge in robotics. Existing methods either limit themselves to highly constrained problems and smaller models to achieve extreme sample efficiency or sacrifice sample efficiency to gain capacity to solve…

Cited by 18SourceScholar
2022

Transferring Dexterous Manipulation from GPU Simulation to a Remote Real-World TriFinger

IROS 2022poster

In-hand manipulation of objects is an important capability to enable robots to carry-out tasks which demand high levels of dexterity. This work presents a robot systems approach to learning dexterous manipulation tasks involving moving objects to arbitrary 6-DoF poses. We show empirical benefits, bo…

Cited by 77SourcecodeScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

Isaac Gym: High Performance GPU Based Physics Simulation For Robot Learning

NeurIPS 2021poster

Isaac Gym offers a high-performance learning platform to train policies for a wide variety of robotics tasks entirely on GPU. Both physics simulation and neural network policy training reside on GPU and communicate by directly passing data from physics buffers to PyTorch tensors without ever going t…

Cited by 969SourcecodeScholar
2020

DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System

ICRA 2020

Teleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usual

Cited by 279SourceScholar
2020

In-Hand Object Pose Tracking via Contact Feedback and GPU-Accelerated Robotic Simulation

ICRA 2020poster

Tracking the pose of an object while it is being held and manipulated by a robot hand is difficult for vision-based methods due to significant occlusions. Prior works have explored using contact feedback and particle filters to localize in-hand objects. However, they have mostly focused on the stati…

Cited by 38SourceScholar
2020

Model-Based Generalization Under Parameter Uncertainty Using Path Integral Control

RA-L 2020

This letter addresses the problem of robot interaction in complex environments where online control and adaptation is necessary. By expanding the sample space in the free energy formulation of path integral control, we derive a natural extension to the path integral control that embeds uncertainty i

Cited by 46SourceScholar
2019

Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience

ICRA 2019poster

We consider the problem of transferring policies to the real world by training on a distribution of simulated scenarios. Rather than manually tuning the randomization of simulations, we adapt the simulation parameter distribution using a few real world roll-outs interleaved with policy training. In…

Cited by 666SourceScholar
2019

ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact

IROS 2019poster

Grasping and manipulating objects is an important human skill. Since most objects are designed to be manipulated by human hands, anthropomorphic hands can enable richer human-robot interaction. Desirable grasps are not only stable, but also functional: they enable post-grasp actions with the object.…

Cited by 132SourceScholar
2019

Learning Latent Space Dynamics for Tactile Servoing

ICRA 2019poster

To achieve a dexterous robotic manipulation, we need to endow our robot with tactile feedback capability, i.e. the ability to drive action based on tactile sensing. In this paper, we specifically address the challenge of tactile servoing, i.e. given the current tactile sensing and a target/goal tact…

Cited by 39SourceScholar
2019

Robust Learning of Tactile Force Estimation through Robot Interaction

ICRA 2019poster

Current methods for estimating force from tactile sensor signals are either inaccurate analytic models or task-specific learned models. In this paper, we explore learning a robust model that maps tactile sensor signals to force. We specifically explore learning a mapping for the SynTouch BioTac sens…

Cited by 71SourceScholar
2018

Domain Randomization and Generative Models for Robotic Grasping

IROS 2018poster

Deep learning-based robotic grasping has made significant progress thanks to algorithmic improvements and increased data availability. However, state-of-the-art models are often trained on as few as hundreds or thousands of unique object instances, and as a result generalization can be a challenge.…

Cited by 194SourceScholar
2018

GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning

CoRL 2018

Most Deep Reinforcement Learning (Deep RL) algorithms require a prohibitively large number of training samples for learning complex tasks. Many recent works on speeding up Deep RL have focused on distributed training and simulation. While distributed training is often done on the GPU, simulation is

2017

SceneNet RGB-D: Can 5M Synthetic Images Beat Generic ImageNet Pre-Training on Indoor Segmentation?

ICCV 2017poster

We introduce SceneNet RGB-D, a dataset providing pixel-perfect ground truth for scene understanding problems such as semantic segmentation, instance segmentation, and object detection. It also provides perfect camera poses and depth data, allowing investigation into geometric computer vision problem…

Cited by 362PDFcodeScholar
2017

SemanticFusion: Dense 3D semantic mapping with convolutional neural networks

ICRA 2017poster

Ever more robust, accurate and detailed mapping using visual sensing has proven to be an enabling factor for mobile robots across a wide variety of applications. For the next level of robot intelligence and intuitive user interaction, maps need to extend beyond geometry and appearance - they need to…

Cited by 830SourceScholar
2016

SceneNet: An annotated model generator for indoor scene understanding

ICRA 2016

We introduce SceneNet, a framework for generating high-quality annotated 3D scenes to aid indoor scene understanding. SceneNet leverages manually-annotated datasets of real world scenes such as NYUv2 to learn statistics about object co-occurrences and their spatial relationships. Using a hierarchica

Cited by 110SourceScholar
2016

Understanding Real World Indoor Scenes With Synthetic Data

CVPR 2016poster

Scene understanding is a prerequisite to many high level tasks for any automated intelligent machine operating in real world environments. Recent attempts with supervised learning have shown promise in this direction but also highlighted the need for enormous quantity of supervised data --- performa…

Cited by 448PDFScholar