← Search

Xinyu Zhan

11 accepted papers

2026

Motion before Action: Diffusing Object Motion As Manipulation Condition

ICRA 2026poster

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations.We propose MBA, a novel module…

2025

AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

CoRL 2025oral

Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection…

Cited by 0SourceScholar
2025

Dense Policy: Bidirectional Autoregressive Learning of Actions

ICCV 2025poster

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for more effective learning methods to unleash the potential of aut…

Cited by 0SourcePDFScholar
2025

Motion Before Action: Diffusing Object Motion as Manipulation Condition

RA-L 2025

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations. We propose MBA (Motion Before

Cited by 16SourceScholar
2024

FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text

AAAI 2024technical

Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virt…

2024

OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion

CVPR 2024poster

We present OAKINK2 a dataset of bimanual object manipulation tasks for complex daily activities. In pursuit of constructing the complex tasks into a structured representation OAKINK2 introduces three level of abstraction to organize the manipulation tasks: Affordance Primitive Task and Complex Task.…

Cited by 18SourcePDFScholar
2023

CHORD: Category-level Hand-held Object Reconstruction via Shape Deformation

ICCV 2023poster

In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of…

Cited by 15PDFcodeScholar
2023

POEM: Reconstructing Hand in a Point Embedded Multi-View Stereo

CVPR 2023poster

Enable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Emb…

2022

ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis

CVPR 2022oral

Estimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can…

Cited by 98PDFcodeScholar
2022

OakInk: A Large-Scale Knowledge Repository for Understanding Hand-Object Interaction

CVPR 2022poster

Learning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases la…

Cited by 97PDFcodeScholar
2021

CPF: Learning a Contact Potential Field To Model the Hand-Object Interaction

ICCV 2021poster

Modeling the hand-object (HO) interaction not only requires estimation of the HO pose, but also pays attention to the contact due to their interaction. Significant progress has been made in estimating hand and object separately with deep learning methods, simultaneous HO pose estimation and contact…

Cited by 141PDFcodeScholar