← Search

Lixin Yang

20 accepted papers

2026

A VAGP-Based Adaptive Kalman Filter for Force Estimation of Robot

RA-L 2026

In human-robot interaction, external force measurement is fundamental to achieving robot compliance control. Parameter identification based on robot dynamics enables external force detection without expensive sensors. However, the unmodeled dynamic errors inherent in real robots pose a challenge to

Cited by 0SourceScholar
2026

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning

ICML 2026poster

Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A major source of failure is their reliance on temporal priors learned from passive video data, which often leads to spat…

Cited by 0SourceScholar
2026

Motion before Action: Diffusing Object Motion As Manipulation Condition

ICRA 2026poster

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations.We propose MBA, a novel module…

2026

TrajBooster: Boosting Humanoid Whole-Body Manipulation Via Trajectory-Centric Learning

ICRA 2026poster

Recent Vision-Language-Action (VLA) models show potential to generalize across embodiments but struggle to quickly align with a new robot’s action space when high-quality demonstrations are scarce, especially for bipedal humanoids. We present TrajBooster, a cross-embodiment framework that leverages …

2025

AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

CoRL 2025oral

Scaling up robotic imitation learning for real-world applications requires efficient and scalable demonstration collection methods. While teleoperation is effective, it depends on costly and inflexible robot platforms. In-the-wild demonstrations offer a promising alternative, but existing collection…

Cited by 0SourceScholar
2025

Dense Policy: Bidirectional Autoregressive Learning of Actions

ICCV 2025poster

Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for more effective learning methods to unleash the potential of aut…

Cited by 0SourcePDFScholar
2025

GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction

CVPR 2025poster

Garments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping s…

Cited by 0SourcePDFScholar
2025

Motion Before Action: Diffusing Object Motion as Manipulation Condition

RA-L 2025

Inferring object motion representations from observations enhances the performance of robotic manipulation tasks. This paper introduces a new paradigm for robot imitation learning that generates action sequences by reasoning about object motion from visual observations. We propose MBA (Motion Before

Cited by 16SourceScholar
2024

FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text

AAAI 2024technical

Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virt…

2024

General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion Process

NeurIPS 2024poster

Articulated object manipulation in real images is a fundamental step in computer and robotic vision tasks. Recently, several image editing methods based on diffusion models have been proposed to manipulate articulated objects according to text prompts. However, these methods often generate weird art…

Cited by 0SourcePDFScholar
2024

OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion

CVPR 2024poster

We present OAKINK2 a dataset of bimanual object manipulation tasks for complex daily activities. In pursuit of constructing the complex tasks into a structured representation OAKINK2 introduces three level of abstraction to organize the manipulation tasks: Affordance Primitive Task and Complex Task.…

Cited by 18SourcePDFScholar
2024

SemGrasp: Semantic Grasp Generation via Language Aligned Discretization

ECCV 2024oral

"Generating natural human grasps necessitates consideration of not just object geometry but also semantic information. Solely depending on object shape for grasp generation confines the applications of prior methods in downstream tasks. This paper presents a novel semantic-based grasp generation met…

2023

CHORD: Category-level Hand-held Object Reconstruction via Shape Deformation

ICCV 2023poster

In daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of…

Cited by 15PDFcodeScholar
2023

POEM: Reconstructing Hand in a Point Embedded Multi-View Stereo

CVPR 2023poster

Enable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Emb…

2022

ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and Synthesis

CVPR 2022oral

Estimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can…

Cited by 98PDFcodeScholar
2022

DART: Articulated Hand Model with Diverse Accessories and Rich Textures

NeurIPS 2022accept

Hand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits it…

2022

ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language Models

NAACL 2022long

Nowadays, pretrained language models (PLMs) have dominated the majority of NLP tasks. While, little research has been conducted on systematically evaluating the language abilities of PLMs. In this paper, we present a large-scale empirical study on general language ability evaluation of PLMs (ElitePL…

2022

OakInk: A Large-Scale Knowledge Repository for Understanding Hand-Object Interaction

CVPR 2022poster

Learning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases la…

Cited by 97PDFcodeScholar
2021

CPF: Learning a Contact Potential Field To Model the Hand-Object Interaction

ICCV 2021poster

Modeling the hand-object (HO) interaction not only requires estimation of the HO pose, but also pays attention to the contact due to their interaction. Significant progress has been made in estimating hand and object separately with deep learning methods, simultaneous HO pose estimation and contact…

Cited by 141PDFcodeScholar
2021

HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation

CVPR 2021poster

Model-based 3D pose and shape estimation methods reconstruct a full 3D mesh for the human body by estimating several parameters. However, learning the abstract parameters is a highly non-linear process and suffers from image-model misalignment, leading to mediocre model performance. In contrast, 3D…

Cited by 469PDFcodeScholar