← Search

Yu-Wei Chao

35 accepted papers

2026

GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks

ICRA 2026poster

Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especially learn-from-demonstration methods, still struggle to handle diverse and dynamic food states, which often results in spillage and reduced reliability.…

2026

GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping

CVPR 2026

We study cross-embodiment 6-DOF robot grasping. Unlike prior works, we require the model not only to generalize to novel objects / scenes but also to novel gripper morphologies and physical grasping processes. Our method extends diffusion model based generative 6-DOF grasping models to condition on

Cited by 0SourcecodeScholar
2026

GraspGen: A Diffusion-Based Framework for 6-DOF Grasping with On-Generator Training

ICRA 2026poster

Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize across different embodiments and in-the-wild settings. We build upon the recent success on modeling the object-centric grasp…

2026

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

CVPR 2026

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: gi

Cited by 0SourcecodeScholar
2025

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference Scoped Exploration

CoRL 2025poster

Hand–object motion-capture (MoCap) repositories provide abundant, contact-rich human demonstrations for scaling dexterous manipulation on robots. Yet demonstration inaccuracy and embodiment gaps between human and robot hands challenge direct policy learning. Existing pipelines adapt a three-stage wo…

Cited by 0SourceScholar
2025

HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction

NeurIPS 2025poster

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGB-D cameras and a HoloLens headset for data collection, avoiding the use of expensive 3D scanners or motion capture systems. We propos…

Cited by 0SourcecodeScholar
2025

Inference-Time Policy Steering Through Human Interactions

ICRA 2025

Generative policies trained with human demonstrations can autonomously accomplish multimodal, longhorizon tasks. However, during inference, humans are often removed from the policy execution loop, limiting the ability to guide a pre-trained policy towards a specific sub-goal or trajectory shape amon

Cited by 37SourcecodeScholar
2025

Latent Action Pretraining from Videos

ICLR 2025poster

We introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators…

Cited by 20SourcePDFScholar
2025

Synthetica: Large Scale Synthetic Data Generation for Robot Perception

IROS 2025

Vision-based object detectors are a crucial basis for robotics applications as they provide valuable information about object localization in the environment. These need to ensure high reliability in different lighting conditions, occlusions, and visual artifacts, all while running in real-time. Col

Cited by 6SourceScholar
2025

TWIN: Two-handed Intelligent Benchmark for Bimanual Manipulation

ICRA 2025

Bimanual manipulation is challenging due to precise spatial and temporal coordination required between two arms. While there exist several real-world bimanual systems, there is a lack of simulated benchmarks with a large task diversity for systematically studying bimanual capabilities across a wide

Cited by 1SourcecodeScholar
2025

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

CoRL 2025poster

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback—a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy…

Cited by 0SourcecodeScholar
2024

Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning

IROS 2024poster

We propose a novel framework for few-shot learning by leveraging large-scale vision-language models such as CLIP [1]. Motivated by unimodal prototypical networks for few-shot learning, we introduce Proto-CLIP which utilizes image prototypes and text prototypes for few-shot learning. Specifically, Pr…

Cited by 8SourcecodeScholar
2024

RVT-2: Learning Precise Manipulation from Few Demonstrations

RSS 2024poster

In this work, we study how to build a robotic system that can solve multiple 3D manipulation tasks given language instructions. To be useful in industrial and household domains, such a system should be capable of learning new tasks with few demonstrations and solving them precisely. Prior works, lik…

2024

SKT-Hang: Hanging Everyday Objects via Object-Agnostic Semantic Keypoint Trajectory Generation

ICRA 2024poster

We study the problem of hanging a wide range of grasped objects on diverse supporting items. Hanging objects is a ubiquitous task that is encountered in numerous aspects of our everyday lives. However, both the objects and supporting items can exhibit substantial variations in their shapes and struc…

Cited by 0SourcecodeScholar
2024

SynH2R: Synthesizing Hand-Object Motions for Learning Human-to-Robot Handovers

ICRA 2024poster

Vision-based human-to-robot handover is an important and challenging task in human-robot interaction. Recent work has attempted to train robot policies by interacting with dynamic virtual humans in simulated environments, where the policies can later be transferred to the real world. However, a majo…

Cited by 20SourceScholar
2023

AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System

RSS 2023poster

Vision-based teleoperation offers the possibility to endow robots with human-level intelligence to physically interact with the environment, while only requiring low-cost camera sensors. However, current vision-based teleoperation systems are designed and engineered towards a particular robot model…

Cited by 114SourcePDFScholar
2023

Learning Human-to-Robot Handovers From Point Clouds

CVPR 2023highlight

We propose the first framework to learn control policies for vision-based human-to-robot handovers, a critical task for human-robot interaction. While research in Embodied AI has made significant progress in training robot agents in simulated environments, interacting with humans remains challenging…

Cited by 50SourcePDFScholar
2023

RVT: Robotic View Transformer for 3D Object Manipulation

CoRL 2023oral

For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adversely affecting scalability. In this work, we propose RVT, a multi-view transform…

Cited by 140SourcecodeScholar
2023

SCONE: A Food Scooping Robot Learning Framework with Active Perception

CoRL 2023poster

Effectively scooping food items poses a substantial challenge for current robotic systems, due to the intricate states and diverse physical properties of food. To address this challenge, we believe in the importance of encoding food items into meaningful representations for effective food scooping.…

Cited by 12SourceScholar
2022

HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers

ICRA 2022poster

We introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics…

Cited by 29SourcecodeScholar
2022

IFOR: Iterative Flow Minimization for Robotic Object Rearrangement

CVPR 2022poster

Accurate object rearrangement from vision is a crucial problem for a wide variety of real-world robotics applications in unstructured environments. We propose IFOR, Iterative Flow Minimization for Robotic Object Rearrangement, an end-to-end method for the challenging problem of object rearrangement…

Cited by 59PDFcodeScholar
2022

Learning Perceptual Concepts by Bootstrapping From Human Queries

RA-L 2022

When robots operate in human environments, it's critical that humans can quickly teach them new <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">concepts:</i> object-centric properties of the environment that they care about (e.g., objects <italic xml

Cited by 17SourceScholar
2022

Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation

CoRL 2022poster

Dexterous robotic hands have the capability to interact with a wide variety of household objects. However, learning robust real world grasping policies for arbitrary objects has proven challenging due to the difficulty of generating high quality training data. In this work, we propose a learning sys…

Cited by 32SourceScholar
2022

Model Predictive Control for Fluid Human-to-Robot Handovers

ICRA 2022poster

Human-robot handover is a fundamental yet challenging task in human-robot interaction and collaboration. Recently, remarkable progressions have been made in human-to-robot handovers of unknown objects by using learning-based grasp generators. However, how to responsively generate smooth motions to t…

Cited by 31SourceScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

Learning to Sit: Synthesizing Human-Chair Interactions via Hierarchical Control

AAAI 2021technical

Recent progress on physics-based character animation has shown impressive breakthroughs on human motion synthesis, through imitating motion capture data via deep reinforcement learning. However, results have mostly been demonstrated on imitating a single distinct motion pattern, and do not generaliz…

Cited by 45SourcePDFScholar
2021

Reactive Human-to-Robot Handovers of Arbitrary Objects

ICRA 2021poster

Human-robot object handovers have been an actively studied area of robotics over the past decade; however, very few techniques and systems have addressed the challenge of handing over diverse objects with arbitrary appearance, size, shape, and deformability. In this paper, we present a vision-based…

Cited by 94SourceScholar
2020

DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System

ICRA 2020

Teleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usual

Cited by 279SourceScholar
2020

Motion Reasoning for Goal-Based Imitation Learning

ICRA 2020poster

We address goal-based imitation learning, where the aim is to output the symbolic goal from a third-person video demonstration. This enables the robot to plan for execution and reproduce the same goal in a completely different environment. The key challenge is that the goal of a video demonstration…

Cited by 20SourceScholar
2018

Rethinking the Faster R-CNN Architecture for Temporal Action Localization

CVPR 2018poster

We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we improve receptive field alignment using a multi-scale architecture that can accom…

Cited by 846SourcePDFScholar
2015

HICO: A Benchmark for Recognizing Human-Object Interactions in Images

ICCV 2015poster

We introduce a new benchmark "Humans Interacting with Common Objects" (HICO) for recognizing human-object interactions (HOI). We demonstrate the key features of HICO: a diverse set of interactions with common object categories, a list of well-defined, sense-based HOI categories, and an exhaustive la…

Cited by 385PDFScholar