← Search

Dieter Fox

177 accepted papers

2026

DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation

ICML 2026poster

We study the problem of functional retargeting: learning dexterous manipulation policies to track object states from human hand-object demonstrations. We focus on long-horizon, bimanual tasks with articulated objects, which are challenging due to large action space, spatiotemporal discontinuities, a…

Cited by 0SourcecodeScholar
2026

GRAPE: Generalizing Robot Policy Via Preference Alignment

ICRA 2026poster

Despite the recent advancements of vision-language-action (VLA) models on a variety of robotics tasks, they suffer from critical issues such as poor generalizability to unseen tasks, due to their reliance on behavior cloning exclusively from successful rollouts. Furthermore, they are typically fine-…

2026

GraspGen: A Diffusion-Based Framework for 6-DOF Grasping with On-Generator Training

ICRA 2026poster

Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize across different embodiments and in-the-wild settings. We build upon the recent success on modeling the object-centric grasp…

2026

MolmoAct: Action Reasoning Models That Can Reason in Space

ICRA 2026poster

Reasoning is essential for purposeful action, yet most robotic foundation models map perception and instructions directly to control, limiting adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models (ARMs), which integrate perception, planning, and control through …

2026

MolmoSpaces: Large-Scale Open Ecosystem for Robot Manipulation and Navigation

RSS 2026poster

Deploying robots at scale demands robustness to the long tail of everyday situations. The countless variations in scene layout, object geometry, and task specifications that characterize real environments are vast and underrepresented in existing robot benchmarks. Measuring this level of generalizat…

Cited by 0SourceScholar
2026

PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies

ICRA 2026poster

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-level reasoning about where and what can be offloaded to vision-language models (VLMs), leaving policies to specialize in h…

2026

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

CVPR 2026

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: gi

Cited by 0SourcecodeScholar
2026

Refinery: Active Fine-Tuning and Deployment-Time Optimization for Contact-Rich Policies

ICRA 2026poster

Simulation-based learning has enabled policies for precise, contact-rich tasks (e.g., robotic assembly) to reach high success rates (~80%) under high levels of observation noise and control error. Although such performance may be sufficient for research applications, it falls short of industry stand…

2026

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

ICRA 2026poster

We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled behavioral and outcome metrics. Existing evaluations often collapse performance into outcome counts, masking differences in execution quality and obscuring fai…

2026

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

RSS 2026poster

General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effective for expert demonstrations, this paradigm scales poorly to large scale real-world robotics datasets where failed and s…

Cited by 0SourceScholar
2026

Uncovering Robot Vulnerabilities through Semantic Potential Fields

ICLR 2026poster

Robot manipulation policies, while central to the promise of physical AI, are highly vulnerable in the presence of external variations in the real world. Diagnosing these vulnerabilities is hindered by two key challenges: (i) the relevant variations to test against are often unknown, and (ii) direct…

Cited by 0SourceScholar
2025

3D-MVP: 3D Multiview Pretraining for Manipulation

CVPR 2025poster

Recent works have shown that visual pretraining on egocentric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics applications require 3D scene understanding. In this work, we pro…

Cited by 0SourcePDFScholar
2025

ACGD: Visual Multitask Policy Learning with Asymmetric Critic Guided Distillation

IROS 2025

We present Asymmetric Critic Guided Distillation, ACGD, a framework for learning multi-task dexterous manipulation policies that can manipulate articulated objects using images as input. ACGD is a scalable student-teacher distillation approach that utilizes behavior cloning to distill multiple exper

Cited by 0SourceScholar
2025

AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

ICLR 2025poster

Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved robots' spatial reasoning and problem-solving abilities, they sti…

2025

Aim My Robot: Precision Local Navigation to Any Object

RA-L 2025

Existing navigation systems mostly consider “success” when the robot reaches within 1 m radius to a goal. This precision is insufficient for emerging applications where a robot needs to be positioned precisely relative to an object for downstream tasks, such as docking, inspection, and manipulation.

Cited by 9SourceScholar
2025

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

CoRL 2025poster

In this work, we unlock new capabilities in robot learning from neural trajectories, synthetic robot data generated from video world models. Our proposed recipe is simple, but powerful: we take the most recent state-of-the-art video generative models (world models), adapt them to the target robot em…

Cited by 0SourcecodeScholar
2025

FORGE: Force-Guided Exploration for Robust Contact-Rich Manipulation Under Uncertainty

RA-L 2025

We present FORGE, a method for sim-to-real transfer of force-aware manipulation policies in the presence of significant pose uncertainty. During simulation-based policy learning, FORGE combines a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">force

Cited by 31SourceScholar
2025

Guiding Long-Horizon Task and Motion Planning with Vision Language Models

ICRA 2025

Vision-Language Models (VLM) can generate plausible high-level plans when prompted with a goal, the context, an image of the scene, and any planning constraints. However, there is no guarantee that the predicted actions are geometrically and kinematically feasible for a particular robot embodiment.

Cited by 68SourcecodeScholar
2025

HAMSTER: Hierarchical Action Models for Open-World Robot Manipulation

ICLR 2025poster

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robotics. One fundamental challenge is the lack of robotic data, which are typically obtained through expensive on-robot ope…

2025

Inference-Time Policy Steering Through Human Interactions

ICRA 2025

Generative policies trained with human demonstrations can autonomously accomplish multimodal, longhorizon tasks. However, during inference, humans are often removed from the policy execution loop, limiting the ability to guide a pre-trained policy towards a specific sub-goal or trajectory shape amon

Cited by 37SourcecodeScholar
2025

Latent Action Pretraining from Videos

ICLR 2025poster

We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators…

Cited by 20SourcePDFScholar
2025

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

CoRL 2025poster

Generative models based on flow matching offer significant potential for learning robot policies, particularly in generating high-dimensional, dexterous behaviors that are conditioned on diverse observations. In this work, we introduce ManiFlow, an advanced flow matching model specifically designed…

Cited by 0SourceScholar
2025

MatchMaker: Automated Asset Generation for Robotic Assembly

ICRA 2025

Robotic assembly remains a significant challenge due to complexities in visual perception, functional grasping, contact-rich manipulation, and performing high-precision tasks. Simulation-based learning and sim-to-real transfer have led to recent success in solving assembly tasks in the presence of o

Cited by 3SourcecodeScholar
2025

OptiGrasp: Optimized Grasp Pose Detection Using RGB Images for Warehouse Picking Robots

IROS 2025

In warehouse environments, robots require robust picking capabilities to manage a wide variety of objects. Effective deployment demands minimal hardware, strong generalization to new products, and resilience in diverse settings. Current methods often rely on depth sensors for structural information,

Cited by 2SourcecodeScholar
2025

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills

NeurIPS 2025poster

Endowing robots with tool design abilities is critical for enabling them to solve complex manipulation tasks that would otherwise be intractable. While recent generative frameworks can automatically synthesize task settings—such as 3D scenes and reward functions—they have not yet addressed the chall…

Cited by 0SourceScholar
2025

SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation

ICML 2025poster

Robotic manipulation systems operating in diverse, dynamic environments must exhibit three critical abilities: multitask interaction, generalization to unseen scenarios, and spatial memory. While significant progress has been made in robotic manipulation, existing approaches often fall short in gene…

2025

SRSA: Skill Retrieval and Adaptation for Robotic Assembly Tasks

ICLR 2025spotlight

Enabling robots to learn novel tasks in a data-efficient manner is a long-standing challenge. Common strategies involve carefully leveraging prior experiences, especially transition data collected on related tasks. Although much progress has been made for general pick-and-place manipulation, far few…

2025

TWIN: Two-handed Intelligent Benchmark for Bimanual Manipulation

ICRA 2025

Bimanual manipulation is challenging due to precise spatial and temporal coordination required between two arms. While there exist several real-world bimanual systems, there is a lack of simulated benchmarks with a large task diversity for systematically studying bimanual capabilities across a wide

Cited by 1SourcecodeScholar
2025

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

CoRL 2025poster

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback—a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy…

Cited by 0SourcecodeScholar
2024

ASID: Active Exploration for System Identification in Robotic Manipulation

ICLR 2024oral

Model-free control strategies such as reinforcement learning have shown the ability to learn control strategies without requiring an accurate model or simulator of the world. While this is appealing due to the lack of modeling requirements, such methods can be sample inefficient, making them impract…

Cited by 14SourcePDFScholar
2024

AutoMate: Specialist and Generalist Assembly Policies over Diverse Geometries

RSS 2024poster

Robotic assembly for high-mixture settings requires adaptivity to diverse parts and poses, which is an open challenge. Meanwhile, in other areas of robotics, large models and sim-to-real have led to tremendous progress. Inspired by such work, we present AutoMate, a learning framework and system that…

Cited by 15SourcePDFScholar
2024

Avoid Everything: Model-Free Collision Avoidance with Expert-Guided Fine-Tuning

CoRL 2024poster

The world is full of clutter. In order to operate effectively in uncontrolled, real world spaces, robots must navigate safely by executing tasks around obstacles while in proximity to hazards. Creating safe movement for robotic manipulators remains a long-standing challenge in robotics, particularly…

Cited by 3SourceScholar
2024

DiMSam: Diffusion Models as Samplers for Task and Motion Planning under Partial Observability

IROS 2024poster

Generative models such as diffusion models, excel at capturing high-dimensional distributions with diverse input modalities, e.g. robot trajectories, but are less effective at multistep constraint reasoning. Task and Motion Planning (TAMP) approaches are suited for planning multi-step autonomous rob…

Cited by 21SourceScholar
2024

DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning

CoRL 2024poster

Running optimization across many parallel seeds leveraging GPU compute [2] have relaxed the need for a good initialization, but this can fail if the problem is highly non-convex as all seeds could get stuck in local minima. One such setting is collision-free motion optimization for robot manipulatio…

Cited by 33SourceScholar
2024

Fast Explicit-Input Assistance for Teleoperation in Clutter

IROS 2024poster

The performance of prediction-based assistance for robot teleoperation degrades in unseen or goal-rich environments due to incorrect or quickly-changing intent inferences. Poor predictions can confuse operators or cause them to change their control input to implicitly signal their goal. We present a…

Cited by 1SourcecodeScholar
2024

IntervenGen: Interventional Data Generation for Robust and Data-Efficient Robot Imitation Learning

IROS 2024poster

Imitation learning is a promising paradigm for training robot control policies, but these policies can suffer from distribution shift, where the conditions at evaluation time differ from those in the training data. A popular approach for increasing policy robustness to distribution shift is interact…

Cited by 8SourceScholar
2024

Learning to Build by Building Your Own Instructions

ECCV 2024poster

"Structural understanding of complex visual objects is an important unsolved component of artificial intelligence. To study this, we develop a new technique for the recently proposed Break-and-Make problem in LTRON where an agent must learn to build a previously unseen LEGO assembly using a single i…

2024

Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

CoRL 2024poster

Large-scale endeavors like RT-1 and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data. However, there is still an opportunity to improve the quality, quantity, and diversity of robot demonstration data. Although vision-language m…

Cited by 39SourcecodeScholar
2024

RVT-2: Learning Precise Manipulation from Few Demonstrations

RSS 2024poster

In this work, we study how to build a robotic system that can solve multiple 3D manipulation tasks given language instructions. To be useful in industrial and household domains, such a system should be capable of learning new tasks with few demonstrations and solving them precisely. Prior works, lik…

2024

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction in Robotics

CoRL 2024poster

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot behavior, VLMs struggle to precisely articulate robot actions usin…

Cited by 53SourcecodeScholar
2024

SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation

CoRL 2024poster

Robot learning has proven to be a general and effective technique for programming manipulators. Imitation learning is able to teach robots solely from human demonstrations but is bottlenecked by the capabilities of the demonstrations. Reinforcement learning uses exploration to discover better behavi…

Cited by 1SourceScholar
2024

SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment

CoRL 2024poster

Imitation learning from human demonstrations is an effective paradigm for robot manipulation, but acquiring large datasets is costly and resource-intensive, especially for long-horizon tasks. To address this issue, we propose SkillGen, an automated system for generating demonstration datasets from a…

Cited by 10SourcecodeScholar
2024

THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation

RSS 2024poster

To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions. Unfortunately, a majority of studies evaluate robot performance in environments closely resembling or even identical to the training setup. We…

2024

URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

RSS 2024poster

Constructing accurate and targeted simulation scenes that are both visually and physically realistic is a problem of significant practical interest in domains ranging from robotics to computer vision. This problem has become even more relevant as researchers wielding large data-hungry learning metho…

Cited by 21SourcePDFScholar
2023

AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System

RSS 2023poster

Vision-based teleoperation offers the possibility to endow robots with human-level intelligence to physically interact with the environment, while only requiring low-cost camera sensors. However, current vision-based teleoperation systems are designed and engineered towards a particular robot model…

Cited by 114SourcePDFScholar
2023

BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects

CVPR 2023poster

We present a near real-time (10Hz) method for 6-DoF tracking of an unknown object from a monocular RGBD video sequence, while simultaneously performing neural 3D reconstruction of the object. Our method works for arbitrary rigid objects, even when visual texture is largely absent. The object is assu…

2023

CabiNet: Scaling Neural Collision Detection for Object Rearrangement with Procedural Scene Generation

ICRA 2023poster

We address the important problem of generalizing robotic rearrangement to clutter without any explicit object models. We first generate over 650K cluttered scenes-orders of magnitude more than prior work-in diverse everyday environments, such as cabinets and shelves. We render synthetic partial poin…

Cited by 27SourcecodeScholar
2023

Constrained Generative Sampling of 6-DoF Grasps

IROS 2023poster

Most state-of-the-art data-driven grasp sampling methods propose stable and collision-free grasps uniformly on the target object. For bin-picking, executing any of those reachable grasps is sufficient. However, for completing specific tasks, such as squeezing out liquid from a bottle, we want the gr…

Cited by 9SourcecodeScholar
2023

CuRobo: Parallelized Collision-Free Robot Motion Generation

ICRA 2023poster

This paper explores the problem of collision-free motion generation for manipulators by formulating it as a global motion optimization problem. We develop a parallel optimization technique to solve this problem and demonstrate its effectiveness on massively parallel GPUs. We show that combining simp…

Cited by 74SourceScholar
2023

DefGraspNets: Grasp Planning on 3D Fields with Graph Neural Nets

ICRA 2023poster

Robotic grasping of 3D deformable objects is critical for real-world applications such as food handling and robotic surgery. Unlike rigid and articulated objects, 3D deformable objects have infinite degrees of freedom. Fully defining their state requires 3D deformation and stress fields, which are e…

Cited by 10SourceScholar
2023

Human-in-the-Loop Task and Motion Planning for Imitation Learning

CoRL 2023poster

Imitation learning from human demonstrations can teach robots complex manipulation skills, but is time-consuming and labor intensive. In contrast, Task and Motion Planning (TAMP) systems are automated and excel at solving long-horizon tasks, but they are difficult to apply to contact-rich tasks. In…

Cited by 21SourcecodeScholar
2023

Imitating Task and Motion Planning with Visuomotor Transformers

CoRL 2023poster

Imitation learning is a powerful tool for training robot manipulation policies, allowing them to learn from expert demonstrations without manual programming or trial-and-error. However, common methods of data collection, such as human supervision, scale poorly, as they are time-consuming and labor-i…

Cited by 56SourcecodeScholar
2023

Impossibly Good Experts and How to Follow Them

ICLR 2023poster

We consider the sequential decision making problem of learning from an expert that has access to more information than the learner. For many problems this extra information will enable the expert to achieve greater long term reward than any policy without this privileged information access. We cal…

Cited by 16SourcePDFScholar
2023

IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality

RSS 2023poster

Robotic assembly is a longstanding challenge, requiring contact-rich interaction and high precision and accuracy. Many applications also require adaptivity to diverse parts, poses, and environments, as well as low cycle times. In other areas of robotics, simulation is a powerful tool to develop algo…

2023

Learning Human-to-Robot Handovers From Point Clouds

CVPR 2023highlight

We propose the first framework to learn control policies for vision-based human-to-robot handovers, a critical task for human-robot interaction. While research in Embodied AI has made significant progress in training robot agents in simulated environments, interacting with humans remains challenging…

Cited by 50SourcePDFScholar
2023

M2T2: Multi-Task Masked Transformer for Object-centric Pick and Place

CoRL 2023poster

With the advent of large language models and large-scale robotic datasets, there has been tremendous progress in high-level decision-making for object manipulation. These generic models are able to interpret complex tasks using language commands, but they often have difficulties generalizing to out-…

Cited by 23SourcecodeScholar
2023

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

CoRL 2023poster

Imitation learning from a large set of human demonstrations has proved to be an effective paradigm for building capable robot agents. However, the demonstrations can be extremely costly and time-consuming to collect. We introduce MimicGen, a system for automatically synthesizing large-scale, rich da…

Cited by 120SourcecodeScholar
2023

NEWTON: Are Large Language Models Capable of Physical Reasoning?

EMNLP 2023long findings

Large Language Models (LLMs), through their contextualized representations, have been empirically proven to encapsulate syntactic, semantic, word sense, and common-sense knowledge. However, there has been limited exploration of their physical reasoning abilities, specifically concerning the crucial…

Cited by 0SourcecodeScholar
2023

ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

ICRA 2023poster

Task planning can require defining myriad domain knowledge about the world in which a robot needs to act. To ameliorate that effort, large language models (LLMs) can be used to score potential next actions during task planning, and even generate action sequences directly, given an instruction in nat…

Cited by 893SourcecodeScholar
2023

RVT: Robotic View Transformer for 3D Object Manipulation

CoRL 2023oral

For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adversely affecting scalability. In this work, we propose RVT, a multi-view transform…

Cited by 140SourcecodeScholar
2023

STOW: Discrete-Frame Segmentation and Tracking of Unseen Objects for Warehouse Picking Robots

CoRL 2023poster

Segmentation and tracking of unseen object instances in discrete frames pose a significant challenge in dynamic industrial robotic contexts, such as distribution warehouses. Here, robots must handle object rearrangements, including shifting, removal, and partial occlusion by new items, and track the…

Cited by 6SourceScholar
2023

Sequence-Based Plan Feasibility Prediction for Efficient Task and Motion Planning

RSS 2023poster

We present a learning-enabled Task and Motion Planning (TAMP) algorithm for solving mobile manipulation problems in environments with many articulated and movable obstacles. Our idea is to bias the search procedure of a traditional TAMP planner with a learned plan feasibility predictor. The core of…

Cited by 30SourcePDFScholar
2023

Shelving, Stacking, Hanging: Relational Pose Diffusion for Multi-modal Rearrangement

CoRL 2023poster

We propose a system for rearranging objects in a scene to achieve a desired object-scene placing relationship, such as a book inserted in an open slot of a bookshelf. The pipeline generalizes to novel geometries, poses, and layouts of both scenes and objects, and is trained from demonstrations to op…

Cited by 46SourcecodeScholar
2023

TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigation

RSS 2023poster

Effective use of camera-based vision systems is essential for robust performance in autonomous off-road driving, particularly in the high-speed regime. Despite success in structured, on-road settings, current end-to-end approaches for scene prediction have yet to be successfully adapted for complex…

Cited by 62SourcePDFScholar
2022

A Bayesian Treatment of Real-to-Sim for Deformable Object Manipulation

RA-L 2022

We consider the problem of inferring simulation parameters such that the behavior of an object in simulation and the real world look similar. This <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">real-to-sim</i> problem is particularly challenging for

Cited by 26SourceScholar
2022

Bayesian Object Models for Robotic Interaction with Differentiable Probabilistic Programming

CoRL 2022poster

A hallmark of human intelligence is the ability to build rich mental models of previously unseen objects from very few interactions. To achieve true, continuous autonomy, robots too must possess this ability. Importantly, to integrate with the probabilistic robotics software stack, such models must…

Cited by 4SourcecodeScholar
2022

Break and Make: Interactive Structural Understanding Using LEGO Bricks

ECCV 2022poster

"Visual understanding of geometric structures with complex spatial relationships is a fundamental component of human intelligence. As children, we learn how to reason about structure not only from observation, but also by interacting with the world around us - by taking things apart and putting them…

2022

Correcting Robot Plans with Natural Language Feedback

RSS 2022poster

When humans design cost or goal specifications for robots, they often produce specifications that are ambiguous, under-specified, or beyond planners’ ability to solve. In these cases, corrections provide a valuable tool for human-in-the-loop robot control. Corrections might take the form of new goal…

Cited by 110SourcePDFScholar
2022

DefGraspSim: Physics-Based Simulation of Grasp Outcomes for 3D Deformable Objects

RA-L 2022

Robotic grasping of 3D deformable objects (e.g., fruits/vegetables, internal organs, bottles/boxes) is critical for real-world applications such as food processing, robotic surgery, and household automation. However, developing grasp strategies for such objects is uniquely challenging. Unlike rigid

Cited by 37SourceScholar
2022

Factory: Fast Contact for Robotic Assembly

RSS 2022poster

Robotic assembly is one of the oldest and most challenging applications of robotics. In other areas of robotics, such as perception and grasping, simulation has rapidly accelerated research progress, particularly when combined with modern deep learning. However, accurately, efficiently, and robustly…

2022

Geometric Fabrics: Generalizing Classical Mechanics to Capture the Physics of Behavior

RA-L 2022

Classical mechanical systems are central to controller design in energy shaping methods of geometric control. However, their expressivity is limited by position-only metrics and the intimate link between metric and geometry. Recent work on Riemannian Motion Policies (RMPs) has shown that shedding th

Cited by 49SourceScholar
2022

HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers

ICRA 2022poster

We introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics…

Cited by 29SourcecodeScholar
2022

IFOR: Iterative Flow Minimization for Robotic Object Rearrangement

CVPR 2022poster

Accurate object rearrangement from vision is a crucial problem for a wide variety of real-world robotics applications in unstructured environments. We propose IFOR, Iterative Flow Minimization for Robotic Object Rearrangement, an end-to-end method for the challenging problem of object rearrangement…

Cited by 59PDFcodeScholar
2022

Learning Perceptual Concepts by Bootstrapping From Human Queries

RA-L 2022

When robots operate in human environments, it's critical that humans can quickly teach them new <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">concepts:</i> object-centric properties of the environment that they care about (e.g., objects <italic xml

Cited by 17SourceScholar
2022

Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation

CoRL 2022poster

Dexterous robotic hands have the capability to interact with a wide variety of household objects. However, learning robust real world grasping policies for arbitrary objects has proven challenging due to the difficulty of generating high quality training data. In this work, we propose a learning sys…

Cited by 32SourceScholar
2022

MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

CoRL 2022poster

We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a region of interest displaying the object in the image and (ii) a CAD model of the observed object. The contributions of thi…

Cited by 157SourcecodeScholar
2022

Model Predictive Control for Fluid Human-to-Robot Handovers

ICRA 2022poster

Human-robot handover is a fundamental yet challenging task in human-robot interaction and collaboration. Recently, remarkable progressions have been made in human-to-robot handovers of unknown objects by using learning-based grasp generators. However, how to responsively generate smooth motions to t…

Cited by 31SourceScholar
2022

Motion Policy Networks

CoRL 2022poster

Collision-free motion generation in unknown environments is a core building block for robot manipulation. Generating such motions is challenging due to multiple objectives; not only should the solutions be optimal, the motion generator itself must be fast enough for real-time performance and reliab…

Cited by 66SourcecodeScholar
2022

Neural Geometric Fabrics: Efficiently Learning High-Dimensional Policies from Demonstration

CoRL 2022poster

Learning dexterous manipulation policies for multi-fingered robots has been a long-standing challenge in robotics. Existing methods either limit themselves to highly constrained problems and smaller models to achieve extreme sample efficiency or sacrifice sample efficiency to gain capacity to solve…

Cited by 18SourceScholar
2022

StructFormer: Learning Spatial Structure for Language-Guided Semantic Rearrangement of Novel Objects

ICRA 2022poster

Geometric organization of objects into semantically meaningful arrangements pervades the built world. As such, assistive robots operating in warehouses, offices, and homes would greatly benefit from the ability to recognize and rearrange objects into these semantically meaningful structures. To be u…

Cited by 104SourceScholar
2022

iCaps: Iterative Category-Level Object Pose and Shape Estimation

RA-L 2022

This letter proposes a category-level 6D object pose and shape estimation approach iCaps, which allows tracking 6D poses of unseen objects in a category and estimating their 3D shapes. We develop a category-level auto-encoder network using depth images as input, where feature embeddings from the aut

Cited by 45SourcecodeScholar
2021

A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution

CoRL 2021poster

Natural language provides an accessible and expressive interface to specify long-term tasks for robotic agents. However, non-experts are likely to specify such tasks with high-level instructions, which abstract over specific robot actions through several layers of abstraction. We propose that key to…

Cited by 151SourcecodeScholar
2021

Alternative Paths Planner (APP) for Provably Fixed-time Manipulation Planning in Semi-structured Environments

ICRA 2021poster

In many applications, including logistics and manufacturing, robot manipulators operate in semi-structured environments alongside humans or other robots. These environments are largely static, but they may contain some movable obstacles that the robot must avoid. Manipulation tasks in these applicat…

Cited by 8SourceScholar
2021

Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes

ICRA 2021poster

Grasping unseen objects in unconstrained, cluttered environments is an essential skill for autonomous robotic manipulation. Despite recent progress in full 6-DoF grasp learning, existing approaches often consist of complex sequential pipelines that possess several potential failure points and run-ti…

Cited by 424SourcecodeScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

DiSECt: A Differentiable Simulation Engine for Autonomous Robotic Cutting

RSS 2021poster

Robotic cutting of soft materials is critical for applications such as food processing; household automation; and surgical manipulation. As in other areas of robotics; simulators can facilitate controller verification; policy learning; and dataset generation. Moreover; differentiable simulators can…

Cited by 113SourcePDFScholar
2021

Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds

CoRL 2021poster

6D robotic grasping beyond top-down bin-picking scenarios is a challenging task. Previous solutions based on 6D grasp synthesis with robot motion planning usually operate in an open-loop setting, which are sensitive to grasp synthesis errors. In this work, we propose a new method for learning closed…

Cited by 55SourcecodeScholar
2021

NeRP: Neural Rearrangement Planning for Unknown Objects

RSS 2021poster

Robots will be expected to manipulate a wide variety of objects in complex and arbitrary ways as they become more widely used in human environments. As such; the rearrangement of objects has been noted to be an important benchmark for AI capabilities in recent years. We propose NeRP (Neural Rearrang…

Cited by 84SourcePDFScholar
2021

Object Rearrangement Using Learned Implicit Collision Functions

ICRA 2021poster

Robotic object rearrangement combines the skills of picking and placing objects. When object models are unavailable, typical collision-checking models may be unable to predict collisions in partial point clouds with occlusions, making generation of collision-free grasping or placement trajectories c…

Cited by 97SourceScholar
2021

Predicting Stable Configurations for Semantic Placement of Novel Objects

CoRL 2021poster

Human environments contain numerous objects configured in a variety of arrangements. Our goal is to enable robots to repose previously unseen objects according to learned semantic relationships in novel environments. We break this problem down into two parts: (1) finding physically valid locations f…

Cited by 54SourcecodeScholar
2021

RGB-D Local Implicit Function for Depth Completion of Transparent Objects

CVPR 2021poster

Majority of the perception methods in robotics require depth information provided by RGB-D cameras. However, standard 3D sensors fail to capture depth of transparent objects due to refraction and absorption of light. In this paper, we introduce a new approach for depth completion of transparent obje…

Cited by 98PDFcodeScholar
2021

RICE: Refining Instance Masks in Cluttered Environments with Graph Neural Networks

CoRL 2021poster

Segmenting unseen object instances in cluttered environments is an important capability that robots need when functioning in unstructured environments. While previous methods have exhibited promising results, they still tend to provide incorrect results in highly cluttered scenes. We postulate that…

Cited by 23SourcecodeScholar
2021

Reactive Human-to-Robot Handovers of Arbitrary Objects

ICRA 2021poster

Human-robot object handovers have been an actively studied area of robotics over the past decade; however, very few techniques and systems have addressed the challenge of handing over diverse objects with arbitrary appearance, size, shape, and deformability. In this paper, we present a vision-based…

Cited by 94SourceScholar
2021

Reactive Long Horizon Task Execution via Visual Skill and Precondition Models

IROS 2021poster

Zero-shot execution of unseen robotic tasks is important to allowing robots to perform a wide variety of tasks in human environments, but collecting the amounts of data necessary to train end-to-end policies in the real-world is often infeasible. We describe an approach for sim-to-real training that…

Cited by 21SourceScholar
2021

Robust Value Iteration for Continuous Control Tasks

RSS 2021poster

When transferring a control policy from simulation to a physical system; this policy needs to be robust to variations in the dynamics to perform well. Commonly; the optimal policy overfits to the approximate model and the corresponding state-distribution. Therefore; the policy fails when transferred…

Cited by 18SourcePDFScholar
2021

SORNet: Spatial Object-Centric Representations for Sequential Manipulation

CoRL 2021oral

Sequential manipulation tasks require a robot to perceive the state of an environment and plan a sequence of actions leading to a desired goal state, where the ability to reason about spatial relationships among object entities from raw sensor inputs is crucial. Prior works relying on explicit state…

Cited by 86SourcecodeScholar
2021

STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation

CoRL 2021oral

Sampling-based model-predictive control (MPC) is a promising tool for feedback control of robots with complex, non-smooth dynamics, and cost functions. However, the computationally demanding nature of sampling-based MPC algorithms has been a key bottleneck in their application to high-dimensional ro…

Cited by 152SourcecodeScholar
2021

Semantic Terrain Classification for Off-Road Autonomous Driving

CoRL 2021poster

Producing dense and accurate traversability maps is crucial for autonomous off-road navigation. In this paper, we focus on the problem of classifying terrains into 4 cost classes (free, low-cost, medium-cost, obstacle) for traversability assessment. This requires a robot to reason about both semanti…

Cited by 100SourceScholar
2021

Sim-to-Real for Robotic Tactile Sensing via Physics-Based Simulation and Learned Latent Projections

ICRA 2021poster

Tactile sensing is critical for robotic grasping and manipulation of objects under visual occlusion. However, in contrast to simulations of robot arms and cameras, current simulations of tactile sensors have limited accuracy, speed, and utility. In this work, we develop an efficient 3D finite elemen…

Cited by 69SourceScholar
2021

Towards Coordinated Robot Motions: End-to-End Learning of Motion Policies on Transform Trees

IROS 2021poster

Generating robot motion that fulfills multiple tasks simultaneously is challenging due to the geometric constraints imposed on the robot. In this paper, we propose to solve multi-task problems through learning structured policies from human demonstrations. Our structured policy is inspired by RMPflo…

Cited by 9SourceScholar
2021

Value Iteration in Continuous Actions, States and Time

ICML 2021spotlight

Classical value iteration approaches are not applicable to environments with continuous states and actions. For such environments the states and actions must be discretized, which leads to an exponential increase in computational complexity. In this paper, we propose continuous fitted value iteratio…

2020

6-DOF Grasping for Target-driven Object Manipulation in Clutter

ICRA 2020poster

Grasping in cluttered environments is a fundamental but challenging robotic skill. It requires both reasoning about unseen object parts and potential collisions with the manipulator. Most existing data-driven approaches avoid this problem by limiting themselves to top-down planar grasps which is ins…

Cited by 263SourcecodeScholar
2020

ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

CVPR 2020poster

We present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks. ALFRED includes long, compositional tasks with non-reversible state changes to shri…

Cited by 922PDFcodeScholar
2020

Camera-to-Robot Pose Estimation from a Single Image

ICRA 2020poster

We present an approach for estimating the pose of an external camera with respect to a robot using a single RGB image of the robot. The image is processed by a deep neural network to detect 2D projections of keypoints (such as joints) associated with the robot. The network is trained entirely on sim…

Cited by 136SourceScholar
2020

Causal Discovery in Physical Systems from Videos

NeurIPS 2020poster

Causal discovery is at the core of human cognition. It enables us to reason about the environment and make counterfactual predictions about unseen scenarios that can vastly differ from our previous experiences. We consider the task of causal discovery from videos in an end-to-end fashion without sup…

Cited by 126SourcePDFScholar
2020

Collaborative Interaction Models for Optimized Human-Robot Teamwork

IROS 2020poster

Effective human-robot collaboration requires informed anticipation. The robot must anticipate the human’s actions, but also react quickly and intuitively when its predictions are wrong. The robot must plan its actions to account for the human’s own plan, with the knowledge that the human’s behavior…

Cited by 22SourceScholar
2020

DeepGMR: Learning Latent Gaussian Mixture Models for Registration

ECCV 2020poster

Point cloud registration is a fundamental problem in 3D computer vision, graphics and robotics. For the last few decades, existing registration algorithms have struggled in situations with large transformations, noise, and time constraints. In this paper, we introduce Deep Gaussian Mixture Registrat…

2020

DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System

ICRA 2020

Teleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usual

Cited by 279SourceScholar
2020

Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning

ICRA 2020poster

Traditional robotic approaches rely on an accurate model of the environment, a detailed description of how to perform the task, and a robust perception system to keep track of the current state. On the other hand, reinforcement learning approaches can operate directly from raw sensory inputs with on…

Cited by 75SourceScholar
2020

IRIS: Implicit Reinforcement without Interaction at Scale for Learning Control from Offline Robot Manipulation Data

ICRA 2020poster

Learning from offline task demonstrations is a problem of great interest in robotics. For simple short-horizon manipulation tasks with modest variation in task instances, offline learning from a small set of demonstrations can produce controllers that successfully solve the task. However, leveraging…

Cited by 146SourceScholar
2020

In-Hand Object Pose Tracking via Contact Feedback and GPU-Accelerated Robotic Simulation

ICRA 2020poster

Tracking the pose of an object while it is being held and manipulated by a robot hand is difficult for vision-based methods due to significant occlusions. Prior works have explored using contact feedback and particle filters to localize in-hand objects. However, they have mostly focused on the stati…

Cited by 38SourceScholar
2020

Inferring the Material Properties of Granular Media for Robotic Tasks

ICRA 2020poster

Granular media (e.g., cereal grains, plastic resin pellets, and pills) are ubiquitous in robotics-integrated industries, such as agriculture, manufacturing, and pharmaceutical development. This prevalence mandates the accurate and efficient simulation of these materials. This work presents a softwar…

Cited by 50SourceScholar
2020

Interpreting and Predicting Tactile Signals via a Physics-Based and Data-Driven Framework

RSS 2020poster

High-density afferents in the human hand have long been regarded as essential for human grasping and manipulation abilities. In contrast, robotic tactile sensors are typically used to provide low-density contact data, such as center-of-pressure and resultant force. Although useful, this data does no…

Cited by 29SourcePDFScholar
2020

LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose Estimation

CVPR 2020poster

Current 6D object pose estimation methods usually require a 3D model for each object. These methods also require additional training in order to incorporate new objects. As a result, they are difficult to scale to a large number of objects and cannot be directly applied to unseen objects. We propose…

Cited by 170PDFScholar
2020

Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

CoRL 2020

Segmenting unseen objects in cluttered scenes is an important skill that robots need to acquire in order to perform tasks in new environments. In this work, we propose a new method for unseen object instance segmentation by learning RGB-D feature embeddings from synthetic data. A metric learning los

Cited by 115SourcePDFScholar
2020

Manipulation Trajectory Optimization with Online Grasp Synthesis and Selection

RSS 2020poster

In robot manipulation, planning the motion of a robot manipulator to grasp an object is a fundamental problem. A manipulation planner needs to generate a trajectory of the manipulator to avoid obstacles in the environment and plan an end-effector pose for grasping. While trajectory planning and gras…

2020

Model-Based Generalization Under Parameter Uncertainty Using Path Integral Control

RA-L 2020

This letter addresses the problem of robot interaction in complex environments where online control and adaptation is necessary. By expanding the sample space in the free energy formulation of path integral control, we derive a natural extension to the path integral control that embeds uncertainty i

Cited by 46SourceScholar
2020

Motion Reasoning for Goal-Based Imitation Learning

ICRA 2020poster

We address goal-based imitation learning, where the aim is to output the symbolic goal from a third-person video demonstration. This enables the robot to plan for execution and reproduce the same goal in a completely different environment. The key challenge is that the goal of a video demonstration…

Cited by 20SourceScholar
2020

Multimodal Trajectory Prediction via Topological Invariance for Navigation at Uncontrolled Intersections

CoRL 2020

We focus on decentralized navigation among multiple non-communicating rational agents at {\em uncontrolled} intersections, i.e., street intersections without traffic signs or signals. Avoiding collisions in such domains relies on the ability of agents to predict each others’ intentions reliably, and

2020

Online BayesSim for Combined Simulator Parameter Inference and Policy Improvement

IROS 2020poster

Recent advancements in Bayesian likelihood-free inference enables a probabilistic treatment for the problem of estimating simulation parameters and their uncertainty given sequences of observations. Domain randomization can be performed much more effectively when a posterior distribution provides th…

Cited by 16SourceScholar
2020

Online Replanning in Belief Space for Partially Observable Task and Motion Problems

ICRA 2020poster

To solve multi-step manipulation tasks in the real world, an autonomous robot must take actions to observe its environment and react to unexpected observations. This may require opening a drawer to observe its contents or moving an object out of the way to examine the space behind it. Upon receiving…

Cited by 145SourcecodeScholar
2020

STReSSD: Sim-To-Real from Sound for Stochastic Dynamics

CoRL 2020

Sound is an information-rich medium that captures dynamic physical events. This work presents STReSSD, a framework that uses sound to bridge the simulation-to-reality gap for stochastic dynamics, demonstrated for the canonical case of a bouncing ball. A physically-motivated noise model is presented

Cited by 0SourcePDFScholar
2020

Self-supervised 6D Object Pose Estimation for Robot Manipulation

ICRA 2020poster

To teach robots skills, it is crucial to obtain data with supervision. Since annotating real world data is time-consuming and expensive, enabling robots to learn in a self- supervised way is important. In this work, we introduce a robot system for self-supervised 6D object pose estimation. Starting…

Cited by 239SourceScholar
2020

Transferable Task Execution from Pixels through Deep Planning Domain Learning

ICRA 2020poster

While robots can learn models to solve many manipulation tasks from raw visual input, they cannot usually use these models to solve new problems. On the other hand, symbolic planning methods such as STRIPS have long been able to solve new problems given only a domain definition and a symbolic goal,…

Cited by 52SourceScholar
2019

BayesSim: Adaptive Domain Randomization Via Probabilistic Inference for Robotics Simulators

RSS 2019poster

We introduce BayesSim, a framework for robotics simulations allowing a full Bayesian treatment for the parameters of the simulator. As simulators become more sophisticated and able to represent the dynamics more accurately, fundamental problems in robotics such as motion planning and perception can…

2019

Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience

ICRA 2019poster

We consider the problem of transferring policies to the real world by training on a distribution of simulated scenarios. Rather than manually tuning the randomization of simulations, we adapt the simulation parameter distribution using a few real world roll-outs interleaved with policy training. In…

Cited by 666SourceScholar
2019

ContactGrasp: Functional Multi-finger Grasp Synthesis from Contact

IROS 2019poster

Grasping and manipulating objects is an important human skill. Since most objects are designed to be manipulated by human hands, anthropomorphic hands can enable richer human-robot interaction. Desirable grasps are not only stable, but also functional: they enable post-grasp actions with the object.…

Cited by 132SourceScholar
2019

EARLY FUSION for Goal Directed Robotic Vision

IROS 2019poster

Building perceptual systems for robotics which perform well under tight computational budgets requires novel architectures which rethink the traditional computer vision pipeline. Modern vision architectures require the agent to build a summary representation of the entire scene, even if most of the…

Cited by 10SourceScholar
2019

Joint Inference of Kinematic and Force Trajectories with Visuo-Tactile Sensing

ICRA 2019poster

To perform complex tasks, robots must be able to interact with and manipulate their surroundings. One of the key challenges in accomplishing this is robust state estimation during physical interactions, where the state involves not only the robot and the object being manipulated, but also the state…

Cited by 38SourceScholar
2019

Learning Latent Space Dynamics for Tactile Servoing

ICRA 2019poster

To achieve a dexterous robotic manipulation, we need to endow our robot with tactile feedback capability, i.e. the ability to drive action based on tactile sensing. In this paper, we specifically address the challenge of tactile servoing, i.e. given the current tactile sensing and a target/goal tact…

Cited by 39SourceScholar
2019

Learning Reactive Motion Policies in Multiple Task Spaces from Human Demonstrations

CoRL 2019

Complex manipulation tasks often require non-trivial and coordinated movements of different parts of a robot. In this work, we address the challenges associated with learning and reproducing the skills required to execute such complex tasks. Specifically, we decompose a task into multiple subtasks a

Cited by 0SourcePDFScholar
2019

Part Segmentation for Highly Accurate Deformable Tracking in Occlusions via Fully Convolutional Neural Networks

ICRA 2019poster

Successfully tracking the human body is an important perceptual challenge for robots that must work around people. Existing methods fall into two broad categories: geometric tracking and direct pose estimation using machine learning. While recent work has shown direct estimation techniques can be qu…

Cited by 5SourceScholar
2019

PoseRBPF: A Rao-Blackwellized Particle Filter for6D Object Pose Estimation

RSS 2019poster

Tracking 6D poses of objects from videos provides rich information to a robot in performing different tasks such as manipulation and navigation. In this work, we formulate the 6D object pose tracking problem in the Rao-Blackwellizedparticle filtering framework, where the 3D rotation and the 3D trans…

Cited by 0SourcePDFScholar
2019

Representing Robot Task Plans as Robust Logical-Dynamical Systems

IROS 2019poster

It is difficult to create robust, reusable, and reactive behaviors for robots that can be easily extended and combined. Frameworks such as Behavior Trees are flexible but difficult to characterize, especially when designing reactions and recovery behaviors to consistently converge to a desired goal…

Cited by 82SourceScholar
2019

Riemannian Motion Policy Fusion through Learnable Lyapunov Function Reshaping

CoRL 2019

RMPflow is a recently proposed policy-fusion framework based on differential geometry. While RMPflow has demonstrated promising performance, it requires the user to provide sensible subtask policies as Riemannian motion policies (RMPs: a motion policy and an importance matrix function), which can be

Cited by 0SourcePDFScholar
2019

Robust Learning of Tactile Force Estimation through Robot Interaction

ICRA 2019poster

Current methods for estimating force from tactile sensor signals are either inaccurate analytic models or task-specific learned models. In this paper, we explore learning a robust model that maps tactile sensor signals to force. We specifically explore learning a mapping for the SynTouch BioTac sens…

Cited by 71SourceScholar
2019

Synthesizing Robot Manipulation Programs from a Single Observed Human Demonstration

IROS 2019poster

Programming by Demonstration (PbD) lets users with little technical background program a wide variety of manipulation tasks for robots, but it should be as intuitive as possible for users while requiring as little time as possible. In this paper, we present a Programming by Demonstration system that…

Cited by 13SourceScholar
2019

The Best of Both Modes: Separately Leveraging RGB and Depth for Unseen Object Instance Segmentation

CoRL 2019

In order to function in unstructured environments, robots need the ability to recognize unseen novel objects. We take a step in this direction by tackling the problem of segmenting unseen object instances in tabletop environments. However, the type of large-scale real-world dataset required for this

Cited by 0SourcePDFScholar
2018

Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects

CoRL 2018

Using synthetic data for training deep neural networks for robotic manipulation holds the promise of an almost unlimited amount of pre-labeled training data, generated safely out of harm’s way. One of the key challenges of synthetic data, to date, has been to bridge the so-called reality gap, so tha

2018

GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning

CoRL 2018

Most Deep Reinforcement Learning (Deep RL) algorithms require a prohibitively large number of training samples for learning complex tasks. Many recent works on speeding up Deep RL have focused on distributed training and simulation. While distributed training is often done on the GPU, simulation is

2018

IQA: Visual Question Answering in Interactive Environments

CVPR 2018poster

We introduce Interactive Question Answering (IQA), the task of answering questions that require an autonomous agent to interact with a dynamic visual environment. IQA presents the agent with a scene and a question, like: “Are there any apples in the fridge?” The agent must navigate around the scene,…

2018

PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes

RSS 2018poster

Estimating the 6D pose of known objects is important for robots to interact with the real world. The problem is challenging due to the variety of objects as well as the complexity of a scene caused by clutter and occlusions between objects. In this work, we introduce PoseCNN, a new Convolutional Neu…

Cited by 2452SourcePDFScholar
2018

Real-time 3D Glint Detection in Remote Eye Tracking Based on Bayesian Inference

ICRA 2018poster

As human gaze provides information on our cognitive states, actions, and intentions, gaze-based interaction has the potential to enable a fluent and natural human-robot collaboration. In this work, we focus on reliable gaze estimation in remote eye tracking based on calibration-free methods. Althoug…

Cited by 14SourceScholar
2018

SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control

ICRA 2018poster

In this work, we present an approach to deep visuomotor control using structured deep dynamics models. Our model, a variant of SE3-Nets, learns a low-dimensional pose embedding for visuomotor control via an encoder-decoder structure. Unlike prior work, our model is structured: given an input scene,…

Cited by 65SourceScholar
2018

Simulating Action Dynamics with Neural Process Networks

ICLR 2018poster

Understanding procedural language requires anticipating the causal effects of actions, even when they are not explicitly stated. In this work, we introduce Neural Process Networks to understand procedural text through (neural) simulation of action dynamics. Our model complements existing memory ar…

Cited by 143SourcePDFScholar
2017

See the Glass Half Full: Reasoning About Liquid Containers, Their Volume and Content

ICCV 2017poster

Humans have rich understanding of liquid containers and their contents; for example, we can effortlessly pour water from a pitcher to a cup. Doing so requires estimating the volume of the cup, approximating the amount of water in the pitcher, and predicting the behavior of water when we tilt the pit…

Cited by 73PDFcodeScholar
2017

Visual Semantic Planning Using Deep Successor Representations

ICCV 2017poster

A crucial capability of real-world intelligent agents is their ability to plan a sequence of actions to achieve their goals in the visual world. In this work, we address the problem of visual semantic planning: the task of predicting a sequence of actions from visual observations that transform a dy…

Cited by 178PDFScholar
2016

Autonomous question answering with mobile robots in human-populated environments

IROS 2016poster

Autonomous mobile robots will soon become ubiquitous in human-populated environments. Besides their typical applications in fetching, delivery, or escorting, such robots present the opportunity to assist human users in their daily tasks by gathering and reporting up-to-date knowledge about the envir…

Cited by 16SourceScholar
2015

Depth-based tracking with physical constraints for robot manipulation

ICRA 2015poster

This work integrates visual and physical constraints to perform real-time depth-only tracking of articulated objects, with a focus on tracking a robot's manipulators and manipulation targets in realistic scenarios. As such, we extend DART, an existing visual articulated object tracker, to additional…

Cited by 83SourceScholar
2015

Designing information gathering robots for human-populated environments

IROS 2015poster

Advances in mobile robotics have enabled robots that can autonomously operate in human-populated environments. Although primary tasks for such robots might be fetching, delivery, or escorting, they present an untapped potential as information gathering agents that can answer questions for the commun…

Cited by 7SourceScholar
2015

DynamicFusion: Reconstruction and Tracking of Non-Rigid Scenes in Real-Time

CVPR 2015poster

We present the first dense SLAM system capable of reconstructing non-rigidly deforming scenes in real-time, by fusing together RGBD scans captured from commodity sensors. Our DynamicFusion approach reconstructs scene geometry whilst simultaneously estimating a dense volumetric 6D motion field that w…

Cited by 1195SourcePDFScholar