← Search

Yu Xiang

54 accepted papers

2026

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes

RSS 2026poster

Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must locate and segment a specific object instance in a cluttered, previously unseen scene. Existing proposal-based approaches…

Cited by 0SourceScholar
2026

Motion Planning with Precedence Specifications Via Augmented Graphs of Convex Sets

ICRA 2026poster

We present an algorithm for planning trajectories that avoid obstacles and satisfy key-door precedence specifi- cations expressed with a fragment of signal temporal logic. Our method includes a novel exact convex partitioning of the obstacle free space that encodes connectivity among convex free spa…

2026

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

RSS 2026poster

General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effective for expert demonstrations, this paradigm scales poorly to large scale real-world robotics datasets where failed and s…

Cited by 0SourceScholar
2025

Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation

IROS 2025

Novel Instance Detection and Segmentation (NIDS) aims at detecting and segmenting novel object instances given a few examples of each instance. We propose a unified, simple, yet effective framework (NIDS-Net) comprising object proposal generation, embedding creation for both instance templates and p

Cited by 13SourcecodeScholar
2025

HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction

NeurIPS 2025poster

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGB-D cameras and a HoloLens headset for data collection, avoiding the use of expensive 3D scanners or motion capture systems. We propos…

Cited by 0SourcecodeScholar
2025

RobotFingerPrint: Unified Gripper Coordinate Space for Multi-Gripper Grasp Synthesis and Transfer

IROS 2025

We introduce a novel grasp representation named the Unified Gripper Coordinate Space (UGCS) for grasp synthesis and grasp transfer. Our representation leverages spherical coordinates to create a shared coordinate space across different robot grippers, enabling it to synthesize and transfer grasps fo

Cited by 1SourceScholar
2025

V-HOP: Visuo-Haptic 6D Object Pose Tracking

RSS 2025poster

Humans naturally integrate vision and haptics for robust object perception during manipulation; losing either modality significantly degrades performance. Inspired by this multisensory integration, prior pose estimation research has attempted to combine visual and haptic/tactile feedback. While thes…

Cited by 2PDFScholar
2024

CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

NeurIPS 2024poster

Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives. These procedures serve as a guiding framework that helps to achieve goals efficiently, whether it is assembling furniture or preparing a recipe. However, the complexity…

Cited by 9SourcePDFScholar
2024

Causal Discovery from Poisson Branching Structural Causal Model Using High-Order Cumulant with Path Analysis

AAAI 2024technical

Count data naturally arise in many fields, such as finance, neuroscience, and epidemiology, and discovering causal structure among count data is a crucial task in various scientific and industrial scenarios. One of the most common characteristics of count data is the inherent branching structure des…

Cited by 2SourcePDFScholar
2024

Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification

AISTATS 2024poster

We present a unified framework called deep dependency networks (DDNs) that combines dependency networks and deep learning architectures for multi-label classification, with a particular emphasis on image and video data. The primary advantage of dependency networks is their ease of training, in contr…

Cited by 2SourcePDFScholar
2024

Grasping Trajectory Optimization with Point Clouds

IROS 2024poster

We introduce a new trajectory optimization method for robotic grasping based on a point-cloud representation of robots and task spaces. In our method, robots are represented by 3D points on their link surfaces. The task space of a robot is represented by a point cloud that can be obtained from depth…

Cited by 2SourceScholar
2024

Mean Shift Mask Transformer for Unseen Object Instance Segmentation

ICRA 2024poster

Segmenting unseen objects from images is a critical perception skill that a robot needs to acquire. In robot manipulation, it can facilitate a robot to grasp and manipulate unseen objects. Mean shift clustering is a widely used method for image segmentation tasks. However, the traditional mean shift…

Cited by 20SourcecodeScholar
2024

MultiGripperGrasp: A Dataset for Robotic Grasping from Parallel Jaw Grippers to Dexterous Hands

IROS 2024poster

We introduce a large-scale dataset named MultiGripperGrasp for robotic grasping. Our dataset contains 30.4M grasps from 11 grippers for 345 objects. These grippers range from two-finger grippers to five-finger grippers, including a human hand. All grasps in the dataset are verified in the robot simu…

Cited by 8SourceScholar
2024

On the Identifiability of Poisson Branching Structural Causal Model Using Probability Generating Function

NeurIPS 2024spotlight

Causal discovery from observational data, especially for count data, is essential across scientific and industrial contexts, such as biology, economics, and network operation maintenance. For this task, most approaches model count data using Bayesian networks or ordinal relations. However, they over…

Cited by 0SourcePDFScholar
2024

Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning

IROS 2024poster

We propose a novel framework for few-shot learning by leveraging large-scale vision-language models such as CLIP [1]. Motivated by unimodal prototypical networks for few-shot learning, we introduce Proto-CLIP which utilizes image prototypes and text prototypes for few-shot learning. Specifically, Pr…

Cited by 8SourcecodeScholar
2024

RISeg: Robot Interactive Object Segmentation via Body Frame-Invariant Features

ICRA 2024poster

In order to successfully perform manipulation tasks in new environments, such as grasping, robots must be proficient in segmenting unseen objects from the background and/or other objects. Previous works perform unseen object instance segmentation (UOIS) by training deep neural networks on large-scal…

Cited by 2SourceScholar
2024

STAF: Pushing the Boundaries of Test-Time Adaptation towards Practical Noise Scenarios

COLING 2024main

Test-time adaptation (TTA) aims to adapt the neural network to the distribution of the target domain using only unlabeled test data. Most previous TTA methods have achieved success under mild conditions, such as considering only a single or multiple independent static domains. However, in real-world…

2024

SceneReplica: Benchmarking Real-World Robot Manipulation by Creating Replicable Scenes

ICRA 2024poster

We present a new reproducible benchmark for evaluating robot manipulation in the real world, specifically focusing on a pick-and-place task. Our benchmark uses the YCB object set, a commonly used dataset in the robotics community, to ensure that our results are comparable to other studies. Additiona…

Cited by 1SourcecodeScholar
2023

Self-Supervised Unseen Object Instance Segmentation via Long-Term Robot Interaction

RSS 2023poster

We introduce a novel robotic system for improving unseen object instance segmentation in the real world by leveraging long-term robot interaction with objects. Previous approaches either grasp or push an object and then obtain the segmentation mask of the grasped or pushed object after one action. I…

Cited by 9SourcePDFScholar
2023

Structural Hawkes Processes for Learning Causal Structure from Discrete-Time Event Sequences

IJCAI 2023poster

Learning causal structure among event types from discrete-time event sequences is a particularly important but challenging task. Existing methods, such as the multivariate Hawkes processes based methods, mostly boil down to learning the so-called Granger causality which assumes that the cause event…

2022

Causal Alignment Based Fault Root Causes Localization for Wireless Network

ICASSP 2022accepted

Localizing fault root causes is challenging but critical for wireless network operation and maintenance. Though supervised methods have shown promising results in training samples, most of the existing approaches assume that the training and the testing samples are independent and identical distribu…

Cited by 0SourceScholar
2022

Few-Shot Single-View 3D Reconstruction with Memory Prior Contrastive Network

ECCV 2022poster

"3D reconstruction of novel categories based on few-shot learning is appealing in real-world applications and attracts increasing research interests. Previous approaches mainly focus on how to design shape prior models for different categories. Their performance on unseen categories is not very comp…

Cited by 20SourcePDFScholar
2022

HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers

ICRA 2022poster

We introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics…

Cited by 29SourcecodeScholar
2022

NeuralGrasps: Learning Implicit Representations for Grasps of Multiple Robotic Hands

CoRL 2022poster

We introduce a neural implicit representation for grasps of objects from multiple robotic hands. Different grasps across multiple robotic hands are encoded into a shared latent space. Each latent vector is learned to decode to the 3D shape of an object and the 3D shape of a robotic hand in a graspin…

Cited by 17SourceScholar
2022

TALISMAN: Targeted Active Learning for Object Detection with Rare Classes and Slices Using Submodular Mutual Information

ECCV 2022poster

"Deep neural networks based object detectors have shown great success in a variety of domains like autonomous vehicles, biomedical imaging, etc. It is known that their success depends on a large amount of data from the domain of interest. While deep models often perform well in terms of overall accu…

2022

iCaps: Iterative Category-Level Object Pose and Shape Estimation

RA-L 2022

This letter proposes a category-level 6D object pose and shape estimation approach iCaps, which allows tracking 6D poses of unseen objects in a category and estimating their 3D shapes. We develop a category-level auto-encoder network using depth images as input, where feature embeddings from the aut

Cited by 45SourcecodeScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds

CoRL 2021poster

6D robotic grasping beyond top-down bin-picking scenarios is a challenging task. Previous solutions based on 6D grasp synthesis with robot motion planning usually operate in an open-loop setting, which are sensitive to grasp synthesis errors. In this work, we propose a new method for learning closed…

Cited by 55SourcecodeScholar
2021

RGB-D Local Implicit Function for Depth Completion of Transparent Objects

CVPR 2021poster

Majority of the perception methods in robotics require depth information provided by RGB-D cameras. However, standard 3D sensors fail to capture depth of transparent objects due to refraction and absorption of light. In this paper, we introduce a new approach for depth completion of transparent obje…

Cited by 98PDFcodeScholar
2021

RICE: Refining Instance Masks in Cluttered Environments with Graph Neural Networks

CoRL 2021poster

Segmenting unseen object instances in cluttered environments is an important capability that robots need when functioning in unstructured environments. While previous methods have exhibited promising results, they still tend to provide incorrect results in highly cluttered scenes. We postulate that…

Cited by 23SourcecodeScholar
2020

LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose Estimation

CVPR 2020poster

Current 6D object pose estimation methods usually require a 3D model for each object. These methods also require additional training in order to incorporate new objects. As a result, they are difficult to scale to a large number of objects and cannot be directly applied to unseen objects. We propose…

Cited by 170PDFScholar
2020

Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation

CoRL 2020

Segmenting unseen objects in cluttered scenes is an important skill that robots need to acquire in order to perform tasks in new environments. In this work, we propose a new method for unseen object instance segmentation by learning RGB-D feature embeddings from synthetic data. A metric learning los

Cited by 115SourcePDFScholar
2020

Manipulation Trajectory Optimization with Online Grasp Synthesis and Selection

RSS 2020poster

In robot manipulation, planning the motion of a robot manipulator to grasp an object is a fundamental problem. A manipulation planner needs to generate a trajectory of the manipulator to avoid obstacles in the environment and plan an end-effector pose for grasping. While trajectory planning and gras…

2020

Self-supervised 6D Object Pose Estimation for Robot Manipulation

ICRA 2020poster

To teach robots skills, it is crucial to obtain data with supervision. Since annotating real world data is time-consuming and expensive, enabling robots to learn in a self- supervised way is important. In this work, we introduce a robot system for self-supervised 6D object pose estimation. Starting…

Cited by 239SourceScholar
2019

PoseRBPF: A Rao-Blackwellized Particle Filter for6D Object Pose Estimation

RSS 2019poster

Tracking 6D poses of objects from videos provides rich information to a robot in performing different tasks such as manipulation and navigation. In this work, we formulate the 6D object pose tracking problem in the Rao-Blackwellizedparticle filtering framework, where the 3D rotation and the 3D trans…

Cited by 0SourcePDFScholar
2019

The Best of Both Modes: Separately Leveraging RGB and Depth for Unseen Object Instance Segmentation

CoRL 2019

In order to function in unstructured environments, robots need the ability to recognize unseen novel objects. We take a step in this direction by tackling the problem of segmenting unseen object instances in tabletop environments. However, the type of large-scale real-world dataset required for this

Cited by 0SourcePDFScholar
2018

Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects

CoRL 2018

Using synthetic data for training deep neural networks for robotic manipulation holds the promise of an almost unlimited amount of pre-labeled training data, generated safely out of harm’s way. One of the key challenges of synthetic data, to date, has been to bridge the so-called reality gap, so tha

2018

Evolutionary Spectra Based on the Multitaper Method with Application To Stationarity Test

ICASSP 2018accepted

In this work, we propose a new inference procedure for understanding non-stationary processes, under the framework of evolutionary spectra developed by Priestley. Among various frameworks of modeling non-stationary processes, the distinguishing feature of the evolutionary spectra is its focus on the…

Cited by 0SourceScholar
2018

PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes

RSS 2018poster

Estimating the 6D pose of known objects is important for robots to interact with the real world. The problem is challenging due to the variety of objects as well as the complexity of a scene caused by clutter and occlusions between objects. In this work, we introduce PoseCNN, a new Convolutional Neu…

Cited by 2452SourcePDFScholar
2016

Deep Metric Learning via Lifted Structured Feature Embedding

CVPR 2016spotlight

Learning the distance metric between pairs of examples is of great importance for learning and visual recognition. With the remarkable success from the state of the art convolutional neural networks, recent works have shown promising results on discriminatively training the networks to learn semanti…

Cited by 2134PDFcodeScholar
2015

A Coarse-to-Fine Model for 3D Pose Estimation and Sub-Category Recognition

CVPR 2015poster

Despite the fact that object detection, 3D pose estimation, and sub-category recognition are highly correlated tasks, they are usually addressed independently from each other because of the huge space of parameters. To jointly model all of these tasks, we propose a coarse-to-fine hierarchical repres…

Cited by 104SourcePDFScholar
2015

Data-Driven 3D Voxel Patterns for Object Category Recognition

CVPR 2015poster

Despite the great progress achieved in recognizing objects as 2D bounding boxes in images, it is still very challenging to detect occluded objects and estimate the 3D properties of multiple objects from a single image. In this paper, we propose a novel object representation, 3D Voxel Pattern (3DVP),…

Cited by 440SourcePDFScholar