← Search

shuran song

112 accepted papers

2026

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

ICLR 2026poster

Recent video generation models demonstrate remarkable ability to capture complex physical interactions and scene evolution over time. To leverage their spatiotemporal priors, robotics works have adapted video models for policy learning but introduce complexity by requiring multiple stages of post-tr…

Cited by 0SourcecodeScholar
2026

DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation

ICML 2026poster

We study the problem of functional retargeting: learning dexterous manipulation policies to track object states from human hand-object demonstrations. We focus on long-horizon, bimanual tasks with articulated objects, which are challenging due to large action space, spatiotemporal discontinuities, a…

Cited by 0SourcecodeScholar
2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

RSS 2026poster

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limi…

Cited by 0SourceScholar
2026

From Prior to Pro: Efficient Skill Mastering via Distribution Contractive RL Finetuning

ICML 2026poster

We introduce Distribution Contractive Reinforcement Learning (DICE-RL), a framework that uses reinforcement learning (RL) as a “distribution contractor” to refine pretrained generative robot policies. DICE-RL turns a pretrained behavior prior into a high-performing “pro” policy by amplifying high-su…

Cited by 0SourcecodeScholar
2026

Geometry-aware 4D Video Generation for Robot Manipulation

ICLR 2026poster

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes, generating videos that are both temporally coherent and geo…

Cited by 0SourcecodeScholar
2026

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations

RSS 2026poster

We present Whole-Body Mobile Manipulation Interface (HoMMI), a data collection and policy learning framework that learns whole-body mobile manipulation directly from robot-free human demonstrations. We augment UMI interfaces with egocentric sensing to capture the global context required for mobile m…

Cited by 0SourceScholar
2026

In-The-Wild Compliant Manipulation with UMI-FT

ICRA 2026poster

Many manipulation tasks require careful force modulation. With insufficient force the task may fail, while excessive force could cause damage. The high cost, bulky size and fragility of commercial force/torque (F/T) sensors have limited large-scale, force-aware policy learning. We introduce UMI-FT, …

2026

SAGE: Scalable Agentic 3D Scene Generation for Embodied AI

CVPR 2026

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines, yielding artifacts and physically invalid scenes. We present

Cited by 0SourcecodeScholar
2026

UMI-On-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

ICRA 2026poster

We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstrations collected with a handheld gripper (UMI) to train generalizable visuomotor policies. A central challenge in transferrin…

2026

UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation

RSS 2026poster

Underwater robotic grasping is difficult due to degraded, highly variable imagery and the expense of collecting diverse underwater demonstrations. We introduce a system that (i) autonomously collects successful underwater grasp demonstrations via a self-supervised data collection pipeline and (ii) t…

Cited by 0SourceScholar
2026

Will People Enjoy a Robot Trainer? a Case Study with Snoopie the Pacerbot

ICRA 2026poster

The physicality of exercise makes the role of athletic trainers unique. Their physical presence allows them to guide a student through a motion, demonstrate an exercise, and give intuitive feedback. Robot quadrupeds are also embodied agents with robust agility and athleticism. In our work, we invest…

2025

A Practical Guide for Incorporating Symmetry in Diffusion Policy

NeurIPS 2025poster

Recently, equivariant neural networks for policy learning have shown promising improvements in sample efficiency and generalization, however, their wide adoption faces substantial barriers due to implementation complexity. Equivariant architectures typically require specialized mathematical formulat…

Cited by 0SourceScholar
2025

Adaptive Compliance Policy: Learning Approximate Compliance for Diffusion Guided Control

ICRA 2025

Compliance plays a crucial role in manipulation, as it balances between the concurrent control of position and force under uncertainties. Yet compliance is often overlooked by today's visuomotor policies that solely focus on position control. This paper introduces Adaptive Compliance Policy (ACP), a

Cited by 57SourcecodeScholar
2025

BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities

CoRL 2025poster

Real-world household tasks present significant challenges for mobile manipulation robots. An analysis of existing robotics benchmarks reveals that successful task performance hinges on three key whole-body control capabilities: bimanual coordination, stable and precise navigation, and extensive end-…

Cited by 0SourcecodeScholar
2025

Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections

NeurIPS 2025poster

We address key challenges in Dataset Aggregation (DAgger) for real-world contact- rich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1)…

Cited by 0SourceScholar
2025

DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation

CoRL 2025oral

We present DexUMI - a data collection and policy learning framework that uses the human hand as the natural interface to transfer dexterous manipulation skills to various robot hands. DexUMI incorporates hardware and software adaptations to minimize the embodiment gap between the human hand and vari…

Cited by 107SourceScholar
2025

Efficient Part-level 3D Object Generation via Dual Volume Packing

NeurIPS 2025poster

Recent progress in 3D object generation has greatly improved both the quality and efficiency. However, most existing methods generate a single mesh with all parts fused together, which limits the ability to edit or manipulate individual parts. A key challenge is that different objects may have a var…

Cited by 0SourcecodeScholar
2025

Language models scale reliably with over-training and on downstream tasks

ICLR 2025poster

Scaling laws are useful guides for derisking expensive training runs, as they predict performance of large models using cheaper, small-scale experiments. However, there remain gaps between current scaling studies and how language models are ultimately trained and evaluated. For instance, scaling is…

2025

Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution

NeurIPS 2025spotlight

Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations from expert trajectories can compound into failure. Common strategies to mitigate this issue involve expanding the training distribution through human-in-the-loop corrections or synthetic data…

Cited by 0SourcecodeScholar
2025

One Demo is Worth a Thousand Trajectories: Action-View Augmentation for Visuomotor Policies

CoRL 2025poster

Visuomotor policies for manipulation have demonstrated remarkable potential in modeling complex robotic behaviors, yet minor alterations in the robot’s initial configuration and unseen obstacles easily lead to out-of-distribution observations. Without extensive data collection effort, these result i…

Cited by 0SourceScholar
2025

Real2Code: Reconstruct Articulated Objects via Code Generation

ICLR 2025poster

We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using image segmentation and shape completion. We represent these object parts with oriented bounding boxes, from which a f…

Cited by 9SourcePDFScholar
2025

Rectified Point Flow: Generic Point Cloud Pose Estimation

NeurIPS 2025spotlight

We present Rectified Point Flow, a unified parameterization that formulates pairwise point cloud registration and multi-part shape assembly as a single conditional generative problem. Given unposed point clouds, our method learns a continuous point-wise velocity field that transports noisy points to…

Cited by 0SourcecodeScholar
2025

Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

CoRL 2025poster

Simulation-based reinforcement learning (RL) has significantly advanced humanoid locomotion tasks, yet direct real-world RL from scratch or starting from pretrained policies remains rare, limiting the full potential of humanoid robots. Real-world training, despite being crucial for overcoming the si…

Cited by 0SourceScholar
2025

Should VLMs be Pre-trained with Image Data?

ICLR 2025poster

Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase effectively unlocks this capability, it is unclear how much of a gain or loss this two-step pipeline gives over VLMs which integrate images earlier int…

Cited by 0SourcePDFScholar
2025

ToddlerBot: Open-Source ML-Compatible Humanoid Platform for Loco-Manipulation

CoRL 2025poster

Learning-based robotics research driven by data demands a new approach to robot hardware design—one that serves as both a platform for policy execution and a tool for embodied data collection. We introduce ToddlerBot, a low-cost, open-source humanoid robot platform designed for robotics and AI resea…

Cited by 0SourceScholar
2025

Vision in Action: Learning Active Perception from Human Demonstrations

CoRL 2025poster

We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic n…

Cited by 0SourceScholar
2024

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

RSS 2024poster

The creation of large, diverse, high-quality robot manipulation datasets is an important stepping stone on the path toward more capable and robust robotic manipulation policies. However, creating such datasets is challenging: collecting robot manipulation data in diverse environments poses logistica…

Cited by 216SourcePDFScholar
2024

DataComp-LM: In search of the next generation of training sets for language models

NeurIPS 2024poster

We introduce DataComp for Language Models, a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl, effective pretraining recipes based on the OpenLM framework, and a broad s…

Cited by 64SourcePDFScholar
2024

Decision Making for Human-in-the-loop Robotic Agents via Uncertainty-Aware Reinforcement Learning

ICRA 2024poster

In a Human-in-the-Loop paradigm, a robotic agent is able to act mostly autonomously in solving a task, but can request help from an external expert when needed. However, knowing when to request such assistance is critical: too few requests can lead to the robot making mistakes, but too many requests…

Cited by 12SourceScholar
2024

DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects

ECCV 2024poster

"Manipulation of elastoplastic objects like dough often involves topological changes such as splitting and merging. The ability to accurately predict these topological changes that a specific action might incur is critical for planning interactions with elastoplastic objects. We present DoughNet, a…

2024

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

CoRL 2024poster

A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a vi…

Cited by 26SourceScholar
2024

EquiBot: SIM(3)-Equivariant Diffusion Policy for Generalizable and Data Efficient Learning

CoRL 2024poster

Building effective imitation learning methods that enable robots to learn from limited data and still generalize across diverse real-world environments is a long-standing problem in robot learning. We propose EquiBot, a robust, data-efficient, and generalizable approach for robot manipulation task l…

Cited by 38SourceScholar
2024

Flow as the Cross-domain Manipulation Interface

CoRL 2024poster

We present Im2Flow2Act, a scalable learning framework that enables robots to acquire real-world manipulation skills without the need of real-world robot training data. The key idea behind Im2Flow2Act is to use object flow as the manipulation interface, bridging domain gaps between different embodime…

Cited by 48SourceScholar
2024

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

CoRL 2024poster

Audio signals provide rich information for the robot interaction and object properties through contact. These information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio…

Cited by 24SourceScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

RIC: Rotate-Inpaint-Complete for Generalizable Scene Reconstruction

ICRA 2024poster

General scene reconstruction refers to the task of estimating the full 3D geometry and texture of a scene containing previously unseen objects. In many practical applications such as AR/VR, autonomous navigation, and robotics, only a single view of the scene may be available, making the scene recons…

Cited by 2SourcecodeScholar
2024

TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning

CoRL 2024poster

Exploiting the promise of recent advances in imitation learning for mobile manipulation will require the collection of large numbers of human-guided demonstrations. This paper proposes an open-source design for an inexpensive, robust, and flexible mobile manipulator that can support arbitrary arms,…

Cited by 6SourceScholar
2024

UMI-on-Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers

CoRL 2024poster

We introduce UMI-on-Legs, a new framework that combines real-world and simulation data for quadruped manipulation systems. We scale task-centric data collection in the real world using a handheld gripper (UMI), providing a cheap way to demonstrate task-relevant manipulation skills without a robot.…

Cited by 45SourceScholar
2024

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

RSS 2024poster

We present Universal Manipulation Interface (UMI) -- a data collection and policy learning framework that allows direct skill transfer from in-the-wild human demonstrations to deployable robot policies. UMI employs hand-held grippers coupled with careful interface design to enable portable, low-cost…

Cited by 235SourcePDFScholar
2023

Bag All You Need: Learning a Generalizable Bagging Strategy for Heterogeneous Objects

IROS 2023poster

We introduce a practical robotics solution for the task of heterogeneous bagging, requiring the placement of multiple rigid and deformable objects into a deformable bag. This is a difficult task as it features complex interactions between multiple highly deformable objects under limited observabilit…

Cited by 19SourceScholar
2023

Cloth Funnels: Canonicalized-Alignment for Multi-Purpose Garment Manipulation

ICRA 2023poster

Automating garment manipulation is challenging due to extremely high variability in object configurations. To reduce this intrinsic variation, we introduce the task of “canonicalized-alignment” that simplifies downstream applications by reducing the possible garment configurations. This task can be…

Cited by 46SourceScholar
2023

CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation

CVPR 2023poster

For robots to be generally useful, they must be able to find arbitrary objects described by people (i.e., be language-driven) even without expensive navigation training on in-domain data (i.e., perform zero-shot inference). We explore these capabilities in a unified setting: language-driven zero-sho…

2023

DataComp: In search of the next generation of multimodal datasets

NeurIPS 2023oral

Multimodal datasets are a critical component in recent breakthroughs such as CLIP, Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the machine learning ecosystem, we introduce Data…

2023

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

RSS 2023poster

This paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot's visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistentl…

Cited by 834SourcePDFScholar
2023

Pick2Place: Task-aware 6DoF Grasp Estimation via Object-Centric Perspective Affordance

ICRA 2023poster

The choice of a grasp plays a critical role in the success of downstream manipulation tasks. Consider a task of placing an object in a cluttered scene; the majority of possible grasps may not be suitable for the desired placement. In this paper, we study the synergy between the picking and placing o…

Cited by 16SourceScholar
2023

REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction

CoRL 2023poster

The ability to detect and analyze failed executions automatically is crucial for an explainable and robust robotic system. Recently, Large Language Models (LLMs) have demonstrated strong reasoning abilities on textual inputs. To leverage the power of LLMs for robot failure explanation, we introduce…

Cited by 136SourcecodeScholar
2023

RoboNinja: Learning an Adaptive Cutting Policy for Multi-Material Objects

RSS 2023poster

We introduce RoboNinja, a learning-based cutting system for multi-material objects (i.e., soft objects with rigid cores such as avocados or mangos). In contrast to prior works using open-loop cutting actions to cut through single-material objects (e.g., slicing a cucumber), RoboNinja aims to remove…

Cited by 30SourcePDFScholar
2023

Structure from Action: Learning Interactions for 3D Articulated Object Structure Discovery

IROS 2023poster

We introduce Structure from Action (SfA), a framework to discover 3D part geometry and joint parameters of unseen articulated objects via a sequence of inferred interactions. Our key insight is that 3D interaction and perception should be considered in conjunction to construct 3D articulated CAD mod…

Cited by 3SourceScholar
2023

TidyBot: Personalized Robot Assistance with Large Language Models

IROS 2023poster

For a robot to personalize physical assistance effectively, it must learn user preferences that can be generally reapplied to future scenarios. In this work, we investigate personalization of household cleanup with robots that can tidy up rooms by picking up objects and putting them away. A key chal…

Cited by 395SourcecodeScholar
2023

Tracking and Reconstructing Hand Object Interactions from Point Cloud Sequences in the Wild

AAAI 2023technical

In this work, we tackle the challenging task of jointly tracking hand object poses and reconstructing their shapes from depth point cloud sequences in the wild, given the initial poses at frame 0. We for the first time propose a point cloud-based hand joint tracking network, HandTrackNet, to estimat…

2022

DextAIRity: Deformable Manipulation Can be a Breeze

RSS 2022poster

This paper introduces DextAIRity, an approach to manipulate deformable objects using active airflow. In contrast to conventional contact-based quasi-static manipulations, DextAIRity allows the system to apply dense forces on out-of-contact surfaces, expands the system's reach range, and provides saf…

Cited by 59SourcePDFScholar
2022

FishGym: A High-Performance Physics-based Simulation Framework for Underwater Robot Learning

ICRA 2022poster

Bionic underwater robots have demonstrated their superiority in many applications. Yet, training their intelligence for a variety of tasks that mimic the behavior of underwater creatures poses a number of challenges in practice, mainly due to lack of a large amount of available training data as well…

Cited by 14SourceScholar
2022

Iterative Residual Policy for Goal-Conditioned Dynamic Manipulation of Deformable Objects

RSS 2022poster

This paper tackles the task of goal-conditioned dynamic manipulation of deformable objects. This task is highly challenging due to its complex dynamics (introduced by object deformation and high-speed action) and strict task requirements (defined by a precise goal specification). To address these ch…

Cited by 90SourcePDFScholar
2022

Learning Pneumatic Non-Prehensile Manipulation With a Mobile Blower

RA-L 2022

We investigate pneumatic non-prehensile manipulation (i.e., blowing) as a means of efficiently moving scattered objects into a target receptacle. Due to the chaotic nature of aerodynamic forces, a blowing controller must i) continually adapt to unexpected changes from its actions, ii) maintain fine-

Cited by 10SourcecodeScholar
2022

Look and Listen: A Multi-Sensory Pouring Network and Dataset for Granular Media from Human Demonstrations

ICRA 2022poster

Humans have the ability to pour various media, both liquid and granular, to desired ends in various containers. We do this by using multiple senses simultaneously in a constant feedback loop to complete a pouring task. Combining multiple sensing modalities, similar to humans, could aid in robotic po…

Cited by 7SourceScholar
2022

Patching open-vocabulary models by interpolating weights

NeurIPS 2022accept

Open-vocabulary models like CLIP achieve high accuracy across many image classification tasks. However, there are still settings where their zero-shot performance is far from optimal. We study model patching, where the goal is to improve accuracy on specific tasks without degrading accuracy on tasks…

2022

Scene Editing as Teleoperation: A Case Study in 6DoF Kit Assembly

IROS 2022poster

Studies in robot teleoperation have been centered around action specifications-from continuous joint control to discrete end-effector pose control. However, these “robot-centric” interfaces often require skilled operators with extensive robotics expertise. To make teleoperation accessible to nonexpe…

Cited by 16SourcecodeScholar
2021

Act the Part: Learning Interaction Strategies for Articulated Object Part Discovery

ICCV 2021poster

People often use physical intuition when manipulating articulated objects, irrespective of object semantics. Motivated by this observation, we identify an important embodied task where an agent must play with objects to recover their parts. To this end, we introduce Act the Part (AtP) to learn how t…

Cited by 50PDFScholar
2021

AdaGrasp: Learning an Adaptive Gripper-Aware Grasping Policy

ICRA 2021poster

This paper aims to improve robots’ versatility and adaptability by allowing them to use a large variety of end- effector tools and quickly adapt to new tools. We propose AdaGrasp, a method to learn a single grasping policy that generalizes to novel grippers. By training on a large collection of grip…

Cited by 52SourcecodeScholar
2021

Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point Clouds

NeurIPS 2021poster

Category-level object pose estimation aims to find 6D object poses of previously unseen object instances from known categories without access to object CAD models. To reduce the huge amount of pose annotations needed for category-level learning, we propose for the first time a self-supervised learni…

2021

SSCNav: Confidence-Aware Semantic Scene Completion for Visual Semantic Navigation

ICRA 2021poster

This paper focuses on visual semantic navigation, the task of producing actions for an active agent to navigate to a specified target object category in an unknown environment. To complete this task, the algorithm should simultaneously locate and navigate to an instance of the category. In compariso…

Cited by 72SourcecodeScholar
2021

Spatial Intention Maps for Multi-Agent Mobile Manipulation

ICRA 2021poster

The ability to communicate intention enables decentralized multi-agent robots to collaborate while performing physical tasks. In this work, we present spatial intention maps, a new intention representation for multi-agent vision-based deep reinforcement learning that improves coordination between de…

Cited by 37SourcecodeScholar
2021

Visual Perspective Taking for Opponent Behavior Modeling

ICRA 2021poster

In order to engage in complex social interaction, humans learn at a young age to infer what others see and cannot see from a different point-of-view, and learn to predict others’ plans and behaviors. These abilities have been mostly lacking in robots, sometimes making them appear awkward and sociall…

Cited by 9SourceScholar
2020

Category-Level Articulated Object Pose Estimation

CVPR 2020oral

This paper addresses the task of category-level pose estimation for articulated objects from a single depth image. We present a novel category-level approach that correctly accommodates object instances previously unseen during training. We introduce Articulation-aware Normalized Coordinate Space Hi…

Cited by 245PDFcodeScholar
2020

Clear Grasp: 3D Shape Estimation of Transparent Objects for Manipulation

ICRA 2020poster

Transparent objects are a common part of everyday life, yet they possess unique visual properties that make them incredibly difficult for standard 3D sensors to produce accurate depth estimates for. In many cases, they often appear as noisy or distorted approximations of the surfaces that lie behind…

Cited by 292SourcecodeScholar
2020

Form2Fit: Learning Shape Priors for Generalizable Assembly from Disassembly

ICRA 2020poster

Is it possible to learn policies for robotic assembly that can generalize to new objects? We explore this idea in the context of the kit assembly task. Since classic methods rely heavily on object pose estimation, they often struggle to generalize to new objects without 3D CAD models or task-specifi…

Cited by 143SourcecodeScholar
2020

Grasping in the Wild: Learning 6DoF Closed-Loop Grasping From Low-Cost Demonstrations

RA-L 2020

Intelligent manipulation benefits from the capacity to flexibly control an end-effector with high degrees of freedom (DoF) and dynamically react to the environment. However, due to the challenges of collecting effective training data and learning efficiently, most grasping algorithms today are limit

Cited by 266SourceScholar
2020

Learning to See before Learning to Act: Visual Pre-training for Manipulation

ICRA 2020poster

Does having visual priors (e.g. the ability to detect objects) facilitate learning to perform vision-based manipulation (e.g. picking up objects)? We study this problem under the framework of transfer learning, where the model is first trained on a passive vision task (i.e., the data distribution do…

Cited by 115SourceScholar
2020

Multitask Learning Strengthens Adversarial Robustness

ECCV 2020poster

Although deep networks achieve strong accuracy on a range of computer vision benchmarks, they remain vulnerable to adversarial attacks, where imperceptible input perturbations fool the network. We present both theoretical and empirical analyses that connect the adversarial robustness of a model to t…

2020

Spatial Action Maps for Mobile Manipulation

RSS 2020poster

Typical end-to-end formulations for learning robotic navigation involve predicting a small set of steering command actions (e.g., step forward, turn left, turn right, etc.) from images of the current state (e.g., a bird's-eye view of a SLAM reconstruction). Instead, we show that it can be advantageo…

2019

DensePhysNet: Learning Dense Physical Object Representations Via Multi-Step Dynamic Interactions

RSS 2019poster

We study the problem of learning physical object representations for robot manipulation. Understanding object physics is critical for successful object manipulation, but also challenging because physical object properties can rarely be inferred from the object's static appearance. In this paper, we…

Cited by 125SourcePDFScholar
2019

Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation

CVPR 2019oral

The goal of this paper is to estimate the 6D pose and dimensions of unseen object instances in an RGB-D image. Contrary to "instance-level" 6D pose estimation tasks, our problem assumes that no exact object CAD models are available during either training or testing time. To handle different and unse…

Cited by 870PDFcodeScholar
2019

TossingBot: Learning to Throw Arbitrary Objects with Residual Physics

RSS 2019poster

We investigate whether a robot arm can learn to pick and throw arbitrary objects into selected boxes quickly and accurately. Throwing has the potential to increase the physical reachability and picking speed of a robot arm. However, precisely throwing arbitrary objects in unstructured settings prese…

Cited by 494SourcePDFScholar
2018

Im2Pano3D: Extrapolating 360° Structure and Semantics Beyond the Field of View

CVPR 2018poster

We present Im2Pano3D, a convolutional neural network that generates a dense prediction of 3D structure and a probability distribution of semantic labels for a full 360 panoramic view of an indoor scene when given only a partial observation ( <=50%) in the form of an RGB-D image. To make this possibl…

2018

Learning Synergies Between Pushing and Grasping with Self-Supervised Deep Reinforcement Learning

IROS 2018poster

Skilled robotic manipulation benefits from complex synergies between non-prehensile (e.g. pushing) and prehensile (e.g. grasping) actions: pushing can help rearrange cluttered objects to make space for arms and fingers; likewise, grasping can help displace objects to make pushing movements more prec…

Cited by 734SourcecodeScholar
2018

Neural Graph Matching Networks for Fewshot 3D Action Recognition

ECCV 2018poster

We propose Neural Graph Matching (NGM) Networks, a novel framework that can learn to recognize a previous unseen 3D action class with only a few examples. We achieve this by leveraging the inherent structure of 3D data through a graphical representation. This allows us to modularize our model and le…

Cited by 132SourcePDFScholar
2018

Robotic Pick-and-Place of Novel Objects in Clutter with Multi-Affordance Grasping and Cross-Domain Image Matching

ICRA 2018poster

This paper presents a robotic pick-and-place system that is capable of grasping and recognizing both known and novel objects in cluttered environments. The key new feature of the system is that it handles a wide range of object categories without needing any task-specific training data for novel obj…

Cited by 848SourcecodeScholar
2017

3DMatch: Learning Local Geometric Descriptors From RGB-D Reconstructions

CVPR 2017oral

Matching local geometric features on real-world depth images is a challenging task due to the noisy, low-resolution, and incomplete nature of 3D scan data. These difficulties limit the performance of current state-of-art methods, which are typically based on histograms over geometric properties. In…

Cited by 1287PDFcodeScholar
2017

Multi-view self-supervised deep learning for 6D pose estimation in the Amazon Picking Challenge

ICRA 2017poster

Robot warehouse automation has attracted significant interest in recent years, perhaps most visibly in the Amazon Picking Challenge (APC) [1]. A fully autonomous warehouse pick-and-place system requires robust vision that reliably recognizes and locates objects amid cluttered environments, self-occl…

Cited by 593SourcecodeScholar
2017

Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks

CVPR 2017poster

Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their representation learning capabilities. One of the bottlenecks in training…

Cited by 329PDFScholar
2017

Semantic Scene Completion From a Single Depth Image

CVPR 2017oral

This paper focuses on semantic scene completion, a task for producing a complete 3D voxel representation of volumetric occupancy and semantic labels for a scene from a single-view depth map observation. Previous work has considered scene completion and semantic labeling of depth maps separately. How…

Cited by 1504PDFcodeScholar
2015

3D ShapeNets: A Deep Representation for Volumetric Shapes

CVPR 2015poster

3D shape is a crucial but heavily underutilized cue in today's computer vision systems, mostly due to the lack of a good generic shape representation. With the recent availability of inexpensive 2.5D depth sensors (e.g. Microsoft Kinect), it is becoming increasingly important to have a powerful 3D s…

Cited by 7454SourcePDFScholar