← Search

Hao-Shu Fang

26 accepted papers

2025

AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

CoRL 2025oral

Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection…

Cited by 0SourceScholar
2025

Dense Policy: Bidirectional Autoregressive Learning of Actions

ICCV 2025poster

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for more effective learning methods to unleash the potential of aut…

Cited by 0SourcePDFScholar
2024

AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild

ICRA 2024poster

While humans can use parts of their arms other than the hands for manipulations like gathering and supporting, whether robots can effectively learn and perform the same type of operations remains relatively unexplored. As these manipulations require joint-level control to regulate the complete poses…

Cited by 40SourcecodeScholar
2024

EyeSight Hand: Design of a Fully-Actuated Dexterous Robot Hand with Integrated Vision-Based Tactile Sensors and Compliant Actuation

IROS 2024poster

In this work, we introduce the EyeSight Hand, a 7 degrees of freedom (DoF) humanoid hand featuring integrated vision-based tactile sensors tailored for enhanced whole-hand manipulation. Additionally, we introduce an actuation scheme centered around quasi-direct drive actuation to achieve human-like…

Cited by 43SourceScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

ICRA 2024poster

A key challenge for robotic manipulation in open domains is how to acquire diverse and generalizable skills for robots. Recent progress in one-shot imitation learning and robotic foundation models have shown promise in transferring trained policies to new tasks based on demonstrations. This feature…

Cited by 86SourcecodeScholar
2024

RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective

IROS 2024poster

Precise robot manipulations require rich spatial information in imitation learning. Image-based policies model object positions from fixed cameras, which are sensitive to camera view changes. Policies utilizing 3D point clouds usually predict keyframes rather than continuous actions, posing difficul…

Cited by 24SourcecodeScholar
2023

Flexible Handover with Real-Time Robust Dynamic Grasp Trajectory Generation

IROS 2023poster

In recent years, there has been a significant effort dedicated to developing efficient, robust, and general human-to-robot handover systems. However, the area of flexible handover in the context of complex and continuous objects' motion remains relatively unexplored. In this work, we propose an appr…

Cited by 8SourceScholar
2023

Target-Referenced Reactive Grasping for Dynamic Objects

CVPR 2023poster

Reactive grasping, which enables the robot to successfully grasp dynamic moving objects, is of great interest in robotics. Current methods mainly focus on the temporal smoothness of the predicted grasp poses but few consider their semantic consistency. Consequently, the predicted grasps are not guar…

Cited by 14SourcePDFScholar
2022

Correlation Field for Boosting 3D Object Detection in Structured Scenes

AAAI 2022technical

Data augmentation is an efficient way to elevate 3D object detection performance. In this paper, we propose a simple but effective online crop-and-paste data augmentation pipeline for structured 3D point cloud scenes, named CorrelaBoost. Observing that 3D objects should have reasonable relative posi…

Cited by 10SourcePDFScholar
2022

Human Trajectory Prediction With Momentary Observation

CVPR 2022poster

Human trajectory prediction task aims to analyze human future movements given their past status, which is a crucial step for many autonomous systems such as self-driving cars and social robots. In real-world scenarios, it is unlikely to obtain sufficiently long observations at all times for predicti…

Cited by 39PDFScholar
2021

DIRV: Dense Interaction Region Voting for End-to-End Human-Object Interaction Detection

AAAI 2021technical

Recent years, human-object interaction (HOI) detection has achieved impressive advances. However, conventional two-stage methods are usually slow in inference. On the other hand, existing one-stage methods mainly focus on the union regions of interactions, which introduce unnecessary visual informat…

2021

Graspness Discovery in Clutters for Fast and Accurate Grasp Detection

ICCV 2021poster

Efficient and robust grasp pose detection is vital for robotic manipulation. For general 6 DoF grasping, conventional methods treat all points in a scene equally and usually adopt uniform sampling to select grasp candidates. However, we discover that ignoring where to grasp greatly harms the speed a…

Cited by 120PDFcodeScholar
2021

RGB Matters: Learning 7-DoF Grasp Poses on Monocular RGBD Images

ICRA 2021poster

General object grasping is an important yet unsolved problem in the field of robotics. Most of the current methods either generate grasp poses with few DoF that fail to cover most of the success grasps, or only take the unstable depth image or point cloud as input which may lead to poor results in s…

Cited by 132SourcecodeScholar
2021

Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and Synthesis

ICCV 2021poster

Multimodal prediction results are essential for trajectory prediction task as there is no single correct answer for the future. Previous frameworks can be divided into three categories: regression, generation and classification frameworks. However, these frameworks have weaknesses in different aspec…

Cited by 89PDFScholar
2020

GraspNet-1Billion: A Large-Scale Benchmark for General Object Grasping

CVPR 2020poster

Object grasping is critical for many applications, which is also a challenging computer vision problem. However, for cluttered scene, current researches suffer from the problems of insufficient training data and the lacking of evaluation benchmarks. In this work, we contribute a large-scale grasp po…

Cited by 650PDFcodeScholar
2019

Cross-Domain Adaptation for Animal Pose Estimation

ICCV 2019oral

In this paper, we are interested in pose estimation of animals. Animals usually exhibit a wide range of variations on poses and there is no available animal pose dataset for training and testing. To address this problem, we build an animal pose dataset to facilitate training and evaluation. Consider…

Cited by 223PDFcodeScholar
2019

CrowdPose: Efficient Crowded Scenes Pose Estimation and a New Benchmark

CVPR 2019oral

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains challenging and inevitable in many scenarios. Moreover, current benchm…

Cited by 699PDFcodeScholar
2019

InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting

ICCV 2019poster

Instance segmentation requires a large number of training samples to achieve satisfactory performance and benefits from proper data augmentation. To enlarge the training set and increase the diversity, previous methods have investigated using data annotation from other domain (e.g. bbox, point) in a…

Cited by 253PDFcodeScholar
2019

Transferable Interactiveness Knowledge for Human-Object Interaction Detection

CVPR 2019poster

Human-Object Interaction (HOI) Detection is an important problem to understand how humans interact with objects. In this paper, we explore Interactiveness Knowledge which indicates whether human and object interact with each other or not. We found that interactiveness knowledge can be learned across…

Cited by 383PDFcodeScholar
2018

Pairwise Body-Part Attention for Recognizing Human-Object Interactions

ECCV 2018poster

In human-object interactions (HOI) recognition, conventional methods consider the human body as a whole and pay a uniform attention to the entire body region. They ignore the fact that normally, human interacts with an object by using some parts of the body. In this paper, we argue that different bo…

Cited by 169SourcePDFScholar
2018

Weakly and Semi Supervised Human Body Part Parsing via Pose-Guided Knowledge Transfer

CVPR 2018poster

Human body part parsing, or human semantic part segmentation, is fundamental to many computer vision tasks. In conventional semantic segmentation methods, the ground truth segmentations are provided, and fully convolutional networks (FCN) are trained in an end-to-end scheme. Although these methods h…