← Search

David Held

80 accepted papers

2026

3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning

ICML 2026poster

We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles. Building on the Deep Latent Particles (DLP) framework, each particle encodes disentangled attributes, including 3D keypoint p…

Cited by 0SourceScholar
2026

Disentangled Point Diffusion for Precise Object Placement

ICRA 2026poster

Recent advances in robotic manipulation have highlighted the effectiveness of learning from demonstration. However, while end-to-end policies excel in expressivity and flexibility, they struggle both in generalizing to novel object geometries and in attaining a high degree of precision. An alternati…

2026

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

RSS 2026poster

We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii…

Cited by 0SourceScholar
2026

Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling

ICLR 2026oral

We introduce Latent Particle World Model (LPWM), a self-supervised object-centric world model scaled to real-world multi-object datasets and applicable in decision-making. LPWM autonomously discovers keypoints, bounding boxes, and object masks directly from video data, enabling it to learn rich scen…

Cited by 0SourcecodeScholar
2025

ArticuBot: Learning Universal Articulated Object Manipulation Policy via Large Scale Simulation

RSS 2025poster

This paper presents ArticuBot, in which a single learned policy enables a robotics system to open diverse categories of unseen articulated objects in the real world. This task has long been challenging for robotics due to the large variations in the geometry, size, and articulation types of such ob…

Cited by 0PDFScholar
2025

Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions

CoRL 2025poster

Robot-assisted dressing has the potential to significantly improve the lives of individuals with mobility impairments. To ensure an effective and comfortable dressing experience, the robot must be able to handle challenging deformable garments, apply appropriate forces, and adapt to limb movements t…

Cited by 0SourceScholar
2025

Geometric Red-Teaming for Robotic Manipulation

CoRL 2025oral

Standard evaluation protocols in robotic manipulation typically assess policy performance over curated, in-distribution test sets, offering limited insight into how systems fail under plausible variation. We introduce a red-teaming framework that probes robustness through object-centric geometr…

Cited by 0SourceScholar
2025

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online

CoRL 2025poster

We introduce a novel History-Aware VErifier (HAVE) to disambiguate uncertain scenarios online by leveraging past interactions. Robots frequently encounter visually ambiguous objects whose manipulation outcomes remain uncertain until physically interacted with. While generative models alone could the…

Cited by 0SourceScholar
2025

Planning from Point Clouds over Continuous Actions for Multi-object Rearrangement

CoRL 2025oral

Multi-object rearrangement is a challenging task that requires robots to reason about a physical 3D scene and the effects of a sequence of actions. While traditional task planning methods are shown to be effective for long-horizon manipulation, they require discretizing the continuous state and acti…

Cited by 0SourceScholar
2025

Real-World Offline Reinforcement Learning from Vision Language Model Feedback

IROS 2025

Offline reinforcement learning can enable policy learning from pre-collected, sub-optimal datasets without online interactions. This makes it ideal for real-world robots and safety-critical scenarios, where collecting online data or expert demonstrations is slow, costly, and risky. However, most exi

Cited by 15SourceScholar
2025

SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting

ICRA 2025

Sim2Real transfer, particularly for manipulation policies relying on RGB images, remains a critical challenge in robotics due to the significant domain shift between syn-thetic and real-world visual data. In this paper, we propose SplatSim, a novel framework that leverages Gaussian Splatting as the

Cited by 66SourcecodeScholar
2024

Deep SE(3)-Equivariant Geometric Reasoning for Precise Placement Tasks

ICLR 2024poster

Many robot manipulation tasks can be framed as geometric reasoning tasks, where an agent must be able to precisely manipulate an object into a position that satisfies the task from a set of initial conditions. Often, task success is defined based on the relationship between two objects - for instanc…

Cited by 14SourcePDFScholar
2024

DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation Learning

NeurIPS 2024spotlight

This paper introduces DiffTORI, which utilizes $\textbf{Diff}$erentiable $\textbf{T}$rajectory $\textbf{O}$ptimization as the policy representation to generate actions for deep $\textbf{R}$einforcement and $\textbf{I}$mitation learning. Trajectory optimization is a powerful and widely used algorithm…

2024

FlowBotHD: History-Aware Diffuser Handling Ambiguities in Articulated Objects Manipulation

CoRL 2024poster

We introduce a novel approach to manipulate articulated objects with ambiguities, such as opening a door, in which multi-modality and occlusions create ambiguities about the opening side and direction. Multi-modality occurs when the method to open a fully closed door (push, pull, slide) is uncertain…

Cited by 0SourcecodeScholar
2024

Force-Constrained Visual Policy: Safe Robot-Assisted Dressing via Multi-Modal Sensing

RA-L 2024

Robot-assisted dressing could profoundly enhance the quality of life of adults with physical disabilities. To achieve this, a robot can benefit from both visual and force sensing. The former enables the robot to ascertain human body pose and garment deformations, while the latter helps maintain safe

Cited by 23SourceScholar
2024

HACMan++: Spatially-Grounded Motion Primitives for Manipulation

RSS 2024poster

Although end-to-end robot learning has shown some success for robot manipulation, the learned policies are often not sufficiently robust to variations in object pose or geometry. To improve the policy generalization, we introduce spatially-grounded parameterized motion primitives in our method HACMa…

2024

Learning Distributional Demonstration Spaces for Task-Specific Cross-Pose Estimation

ICRA 2024poster

Relative placement tasks are an important category of tasks in which one object needs to be placed in a desired pose relative to another object. Previous work has shown success in learning relative placement tasks from just a small number of demonstrations when using relational reasoning networks wi…

Cited by 1SourceScholar
2024

Learning Generalizable Tool-use Skills through Trajectory Generation

IROS 2024poster

Autonomous systems that efficiently utilize tools can assist humans in completing many common tasks such as cooking and cleaning. However, current systems fall short of matching human-level of intelligence in terms of adapting to novel tools. Prior works based on affordance often make strong assumpt…

Cited by 3SourceScholar
2024

Modeling Drivers’ Situational Awareness from Eye Gaze for Driving Assistance

CoRL 2024poster

Intelligent driving assistance can alert drivers to objects in their environment; however, such systems require a model of drivers' situational awareness (SA) (what aspects of the scene they are already aware of) to avoid unnecessary alerts. Moreover, collecting the data to train such an SA model i…

Cited by 1SourceScholar
2024

Object Importance Estimation Using Counterfactual Reasoning for Intelligent Driving

RA-L 2024

The ability to identify important objects in a complex and dynamic driving environment is essential for autonomous driving agents to make safe and efficient driving decisions. It also helps assistive driving systems decide when to alert drivers. We tackle object importance estimation in a data-drive

Cited by 5SourceScholar
2024

RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

ICML 2024poster

Reward engineering has long been a challenge in Reinforcement Learning (RL) research, as it often requires extensive human effort and iterative processes of trial-and-error to design effective reward functions. In this paper, we propose RL-VLM-F, a method that automatically generates reward function…

2024

Reinforcement Learning in a Safety-Embedded MDP with Trajectory Optimization

ICRA 2024poster

Safe Reinforcement Learning (RL) plays an important role in applying RL algorithms to safety-critical real-world applications, addressing the trade-off between maximizing rewards and adhering to safety constraints. This work introduces a novel approach that combines RL with trajectory optimization t…

Cited by 1SourceScholar
2024

RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

ICML 2024poster

We present RoboGen, a generative robotic agent that automatically learns diverse robotic skills at scale via generative simulation. RoboGen leverages the latest advancements in foundation and generative models. Instead of directly adapting these models to produce policies or low-level actions, we ad…

Cited by 88SourcePDFScholar
2023

Active Velocity Estimation using Light Curtains via Self-Supervised Multi-Armed Bandits

RSS 2023poster

To navigate in an environment safely and autonomously, robots must accurately estimate where obstacles are and how they move. Instead of using expensive traditional 3D sensors, we explore the use of a much cheaper, faster, and higher resolution alternative: programmable light curtains. Light curtain…

Cited by 1SourcePDFScholar
2023

AutoBag: Learning to Open Plastic Bags and Insert Objects

ICRA 2023poster

Thin plastic bags are ubiquitous in retail stores, healthcare, food handling, recycling, homes, and school lunchrooms. They are challenging both for perception (due to specularities and occlusions) and for manipulation (due to the dynamics of their 3D deformable structure). We formulate the task of…

Cited by 44SourceScholar
2023

Bagging by Learning to Singulate Layers Using Interactive Perception

IROS 2023poster

Many fabric handling and 2D deformable material tasks in homes and industries require singulating layers of material such as opening a bag or arranging garments for sewing. In contrast to methods requiring specialized sensing or end effectors, we use only visual observations with ordinary parallel j…

Cited by 12SourceScholar
2023

EDO-Net: Learning Elastic Properties of Deformable Objects from Graph Dynamics

ICRA 2023poster

We study the problem of learning graph dynamics of deformable objects that generalizes to unknown physical properties. Our key insight is to leverage a latent representation of elastic physical properties of cloth-like deformable objects that can be extracted, for example, from a pulling interaction…

Cited by 27SourceScholar
2023

Elastic Context: Encoding Elasticity for Data-driven Models of Textiles Elastic Context: Encoding Elasticity for Data-driven Models of Textiles

ICRA 2023poster

Physical interaction with textiles, such as assistive dressing or household tasks, requires advanced dexterous skills. The complexity of textile behavior during stretching and pulling is influenced by the material properties of the yarn and by the textile's construction technique, which are often un…

Cited by 10SourceScholar
2023

FlowBot++: Learning Generalized Articulated Objects Manipulation via Articulation Projection

CoRL 2023poster

Understanding and manipulating articulated objects, such as doors and drawers, is crucial for robots operating in human environments. We wish to develop a system that can learn to articulate novel objects with no prior interaction, after training on other articulated objects. Previous approaches fo…

Cited by 35SourceScholar
2023

HACMan: Learning Hybrid Actor-Critic Maps for 6D Non-Prehensile Manipulation

CoRL 2023oral

Manipulating objects without grasping them is an essential component of human dexterity, referred to as non-prehensile manipulation. Non-prehensile manipulation may enable more complex interactions with the objects, but also presents challenges in reasoning about gripper-object interactions. In this…

Cited by 22SourcecodeScholar
2023

One Policy to Dress Them All: Learning to Dress People with Diverse Poses and Garments

RSS 2023poster

Robot-assisted dressing could benefit the lives of many people such as older adults and individuals with disabilities. Despite such potential, robot-assisted dressing remains a challenging task for robotics as it involves complex manipulation of deformable cloth in 3D space. Many prior works aim to…

Cited by 21SourcePDFScholar
2023

Point Cloud Forecasting as a Proxy for 4D Occupancy Forecasting

CVPR 2023poster

Predicting how the world can evolve in the future is crucial for motion planning in autonomous systems. Classical methods are limited because they rely on costly human annotations in the form of semantic class labels, bounding boxes, and tracks or HD maps of cities to plan their motion -- and thus a…

2022

DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools

ICLR 2022poster

We consider the problem of sequential robotic manipulation of deformable objects using tools. Previous works have shown that differentiable physics simulators provide gradients to the environment state and help trajectory optimization to converge orders of magnitude faster than model-free reinforcem…

Cited by 64SourcePDFScholar
2022

Differentiable Raycasting for Self-Supervised Occupancy Forecasting

ECCV 2022poster

"Motion planning for safe autonomous driving requires learning how the environment around an ego-vehicle evolves with time. Ego-centric perception of driveable regions in a scene not only changes with the motion of actors in the environment, but also with the movement of the ego-vehicle itself. Self…

2022

Learning to Singulate Layers of Cloth using Tactile Feedback

IROS 2022poster

Robotic manipulation of cloth has applications ranging from fabrics manufacturing to handling blankets and laundry. Cloth manipulation is challenging for robots largely due to their high degrees of freedom, complex dynamics, and severe self-occlusions when in folded or crumpled configurations. Prior…

Cited by 22SourceScholar
2022

Planning with Spatial-Temporal Abstraction from Point Clouds for Deformable Object Manipulation

CoRL 2022poster

Effective planning of long-horizon deformable object manipulation requires suitable abstractions at both the spatial and temporal levels. Previous methods typically either focus on short-horizon tasks or make strong assumptions that full-state information is available, which prevents their use on de…

Cited by 39SourceScholar
2022

Self-supervised Transparent Liquid Segmentation for Robotic Pouring

ICRA 2022poster

Liquid state estimation is important for robotics tasks such as pouring; however, estimating the state of transparent liquids is a challenging problem. We propose a novel segmentation pipeline that can segment transparent liquids such as water from a static, RGB image without requiring any manual an…

Cited by 23SourcecodeScholar
2022

TAX-Pose: Task-Specific Cross-Pose Estimation for Robot Manipulation

CoRL 2022poster

How do we imbue robots with the ability to efficiently manipulate unseen objects and transfer relevant skills based on demonstrations? End-to-end learning methods often fail to generalize to novel objects or unseen configurations. Instead, we focus on the task-specific pose relationship between rele…

Cited by 61SourceScholar
2022

ToolFlowNet: Robotic Manipulation with Tools via Predicting Tool Flow from Point Clouds

CoRL 2022poster

Point clouds are a widely available and canonical data modality which convey the 3D geometry of a scene. Despite significant progress in classification and segmentation from point clouds, policy learning from such a modality remains challenging, and most prior works in imitation learning focus on le…

Cited by 59SourceScholar
2022

Visual Haptic Reasoning: Estimating Contact Forces by Observing Deformable Object Interactions

RA-L 2022

Robotic manipulation of highly deformable cloth presents a promising opportunity to assist people with several daily tasks, such as washing dishes; folding laundry; or dressing, bathing, and hygiene assistance for individuals with severe motor impairments. In this letter, we introduce a formulation

Cited by 23SourceScholar
2021

Active Safety Envelopes using Light Curtains with Probabilistic Guarantees

RSS 2021poster

To safely navigate unknown environments; robots must accurately perceive dynamic obstacles. Instead of directly measuring the scene depth with a LiDAR sensor; we explore the use of a much cheaper and higher resolution sensor: programmable light curtains. Light curtains are controllable depth sensors…

Cited by 7SourcePDFScholar
2021

Exploiting & Refining Depth Distributions With Triangulation Light Curtains

CVPR 2021poster

Active sensing through the use of Adaptive Depth Sensors is a nascent field, with potential in areas such as Advanced driver-assistance systems (ADAS). They do however require dynamically driving a laser / light-source to a specific location to capture information, with one such class of sensor bein…

Cited by 9PDFScholar
2021

FabricFlowNet: Bimanual Cloth Manipulation with a Flow-based Policy

CoRL 2021poster

We address the problem of goal-directed cloth manipulation, a challenging task due to the deformability of cloth. Our insight is that optical flow, a technique normally used for motion estimation in video, can also provide an effective representation for corresponding cloth poses across observation…

Cited by 99SourcecodeScholar
2021

RB2: Robotic Manipulation Benchmarking with a Twist

NeurIPS 2021poster

Benchmarks offer a scientific way to compare algorithms using objective performance metrics. Good benchmarks have two features: (a) they should be widely useful for many research groups; (b) and they should produce reproducible findings. In robotic manipulation research, there is a trade-off between…

Cited by 25SourceScholar
2021

Safe Local Motion Planning With Self-Supervised Freespace Forecasting

CVPR 2021poster

Safe local motion planning for autonomous driving in dynamic environments requires forecasting how the scene evolves. Practical autonomy stacks adopt a semantic object-centric representation of a dynamic scene and build object detection, tracking, and prediction modules to solve forecasting. However…

Cited by 94PDFcodeScholar
2020

3D Multi-Object Tracking: A Baseline and New Evaluation Metrics

IROS 2020poster

3D multi-object tracking (MOT) is an essential component for many applications such as autonomous driving and assistive robotics. Recent work on 3D MOT focuses on developing accurate systems giving less attention to practical considerations such as computational cost and system complexity. In contra…

Cited by 543SourcecodeScholar
2020

Active Perception using Light Curtains for Autonomous Driving

ECCV 2020poster

Most real-world 3D sensors such as LiDARs are passive, meaning that they sense the entire environment, while being decoupled from the recognition system that processes the sensor data. In this work, we propose a method for 3D object recognition using light curtains, a resource-efficient active senso…

Cited by 14SourcePDFScholar
2020

Cloth Region Segmentation for Robust Grasp Selection

IROS 2020poster

Cloth detection and manipulation is a common task in domestic and industrial settings, yet such tasks remain a challenge for robots due to cloth deformability. Furthermore, in many cloth-related tasks like laundry folding and bed making, it is crucial to manipulate specific regions like edges and co…

Cited by 58SourcecodeScholar
2020

Learning Orientation Distributions for Object Pose Estimation

IROS 2020poster

For robots to operate robustly in the real world, they should be aware of their uncertainty. However, most methods for object pose estimation return a single point estimate of the object's pose. In this work, we propose two learned methods for estimating a distribution over an object's orientation.…

Cited by 23SourcecodeScholar
2020

Multi-Modal Transfer Learning for Grasping Transparent and Specular Objects

RA-L 2020

State-of-the-art object grasping methods rely on depth sensing to plan robust grasps, but commercially available depth sensors fail to detect transparent and specular objects. To improve grasping performance on such objects, we introduce a method for learning a multi-modal perception model by bootst

Cited by 38SourceScholar
2020

ROLL: Visual Self-Supervised Reinforcement Learning with Object Reasoning

CoRL 2020

Current image-based reinforcement learning (RL) algorithms typically operate on the whole image without performing object-level reasoning. This leads to inefficient goal sampling and ineffective reward functions. In this paper, we improve upon previous visual self-supervised RL by incorporating obje

2020

SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation

CoRL 2020

Manipulating deformable objects has long been a challenge in robotics due to its high dimensional state representation and complex dynamics. Recent success in deep reinforcement learning provides a promising direction for learning to manipulate deformable objects with data driven methods. However, e

2020

What You See is What You Get: Exploiting Visibility for 3D Object Detection

CVPR 2020oral

Recent advances in 3D sensing have created unique challenges for computer vision. One fundamental challenge is finding a good representation for 3D sensor data. Most popular representations (such as PointNet) are proposed in the context of processing truly 3D data (e.g. points sampled from mesh mode…

Cited by 151PDFcodeScholar
2019

Adaptive Auxiliary Task Weighting for Reinforcement Learning

NeurIPS 2019poster

Reinforcement learning is known to be sample inefficient, preventing its application to many real-world problems, especially with high dimensional observations like images. Transferring knowledge from other auxiliary tasks is a powerful tool for improving the learning efficiency. However, the usage…

2019

Combining Deep Learning and Verification for Precise Object Instance Detection

CoRL 2019

Deep learning based object detectors often report false positives with very high confidence. Although they optimize generic detection performance, such as mean average precision (mAP), they are not designed for robustness or verifiability. We argue that, if a high confidence detection is made by a r

2018

Automatic Goal Generation for Reinforcement Learning Agents

ICML 2018oral

Reinforcement learning (RL) is a powerful technique to train an agent to perform a task; however, an agent that is trained using RL is only capable of achieving the single task that is specified via its reward function. Such an approach does not scale well to settings in which an agent needs to perf…

Cited by 530SourcePDFScholar
2017

Reverse Curriculum Generation for Reinforcement Learning

CoRL 2017

Many relevant tasks require an agent to reach a certain state, or to manipulate objects into a desired configuration. For example, we might want a robot to align and assemble a gear onto an axle or insert and turn a key in a lock. These goal-oriented tasks present a considerable challenge for reinfo

Cited by 0SourcePDFScholar
2016

A Probabilistic Framework for Real-time 3D Segmentation using Spatial, Temporal, and Semantic Cues

RSS 2016poster

In order to track dynamic objects in a robot’s environment, one must first segment the scene into a collection of separate objects. Most real-time robotic vision systems today rely on simple spatial relations to segment the scene into separate objects. However, such methods fail under a variety of…

Cited by 54SourcePDFScholar