← Search

Yanjie Ze

21 accepted papers

2026

Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking

ICRA 2026poster

Humanoid motion tracking policies are central to building teleoperation pipelines and hierarchical controllers, yet they face a fundamental challenge: the embodiment gap between humans and humanoid robots. Current approaches address this gap by retargeting human motion data to humanoid embodiments a…

2026

TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System

ICRA 2026poster

Large-scale data has driven breakthroughs in robotics, from language models to vision-language-action models in bimanual manipulation. However, humanoid robotics lacks equally effective data collection frameworks. Existing humanoid teleoperation systems either use decoupled control or depend on expe…

2025

4D Visual Pre-training for Robot Learning

ICCV 2025poster

General visual representations learned from web-scale datasets for robotics have achieved great success in recent years, enabling data-efficient robot learning on manipulation tasks; yet these pre-trained representations are mostly on 2D images, neglecting the inherent 3D nature of the world. Howeve…

2025

BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities

CoRL 2025poster

Real-world household tasks present significant challenges for mobile manipulation robots. An analysis of existing robotics benchmarks reveals that successful task performance hinges on three key whole-body control capabilities: bimanual coordination, stable and precise navigation, and extensive end-…

Cited by 0SourcecodeScholar
2025

Catch It! Learning to Catch in Flight with Mobile Dexterous Hands

ICRA 2025

Catching objects in flight (i.e., thrown objects) is a common daily skill for humans, yet it presents a significant challenge for robots. This task requires a robot with agile and accurate motion, a large spatial workspace, and the ability to interact with diverse objects. In this paper, we build a

Cited by 27SourcecodeScholar
2025

Generalizable Humanoid Manipulation with 3D Diffusion Policies

IROS 2025

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the difficulty of acquiring generalizable skills and the expensivenes

Cited by 40SourcecodeScholar
2025

Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies

IROS 2025

Reinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, such as low-pass filters and smoothness rewards, are often employed to develop po

Cited by 49SourcecodeScholar
2025

TWIST: Teleoperated Whole-Body Imitation System

CoRL 2025poster

Teleoperating humanoid robots in a whole-body manner marks a fundamental step toward developing general-purpose robotic intelligence, with human motion providing an ideal interface for controlling all degrees of freedom. Yet, most current humanoid teleoperation systems fall short of enabling coordin…

Cited by 0SourceScholar
2025

X-Capture: An Open-Source Portable Device for Multi-Sensory Learning

ICCV 2025poster

Understanding objects through multiple sensory modalities is fundamental to human perception, enabling cross-sensory integration and richer comprehension. For AI and robotic systems to replicate this ability, access to diverse, high-quality multi-sensory data is critical. Existing datasets are often…

Cited by 0SourcePDFScholar
2024

3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

RSS 2024poster

Imitation learning provides an efficient way to teach robots dexterous skills; however, learning complex skills robustly and generalizablely usually consumes large amounts of human demonstrations. To tackle this challenging problem, we present 3D Diffusion Policy (DP3), a novel visual imitation lear…

2024

Diffusion Reward: Learning Rewards via Conditional Video Diffusion

ECCV 2024poster

"Learning rewards from expert videos offers an affordable and effective solution to specify the intended behaviors for reinforcement learning (RL) tasks. In this work, we propose , a novel framework that learns rewards from expert videos via conditional video diffusion models for solving complex vis…

2024

DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization

ICLR 2024spotlight

Visual reinforcement learning (RL) has shown promise in continuous control tasks. Despite its progress, current algorithms are still unsatisfactory in virtually every aspect of the performance such as sample efficiency, asymptotic performance, and their robustness to the choice of random seeds. In t…

2024

Learning Visual Quadrupedal Loco-Manipulation from Demonstrations

IROS 2024poster

Quadruped robots are progressively being integrated into human environments. Despite the growing locomotion capabilities of quadrupedal robots, their interaction with objects in realistic scenes is still limited. While additional robotic arms on quadrupedal robots enable manipulating objects, they a…

Cited by 17SourceScholar
2024

Unleashing the Power of Pre-trained Language Models for Offline Reinforcement Learning

ICLR 2024poster

Offline reinforcement learning (RL) aims to find a near-optimal policy using pre-collected datasets. Given recent advances in Large Language Models (LLMs) and their few-shot learning prowess, this paper introduces $\textbf{La}$nguage Models for $\textbf{Mo}$tion Control ($\textbf{LaMo}$), a general…

2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

IJCAI 2023poster

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating with others, yet such privacy concern has not been considered in existing works in MARL. We propose the differentially…

2023

GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields

CoRL 2023oral

It is a long-standing problem in robotics to develop agents capable of executing diverse manipulation tasks from visual observations in unstructured real-world environments. To achieve this goal, the robot will need to have a comprehensive understanding of the 3D structure and semantics of the scen…

Cited by 88SourcecodeScholar
2023

H-InDex: Visual Reinforcement Learning with Hand-Informed Representations for Dexterous Manipulation

NeurIPS 2023poster

Human hands possess remarkable dexterity and have long served as a source of inspiration for robotic manipulation. In this work, we propose a human $\textbf{H}$and-$\textbf{In}$formed visual representation learning framework to solve difficult $\textbf{Dex}$terous manipulation tasks ($\textbf{H-InDe…

Cited by 21SourcePDFScholar
2023

On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline

ICML 2023poster

In this paper, we examine the effectiveness of pre-training for visuo-motor control tasks. We revisit a simple Learning-from-Scratch (LfS) baseline that incorporates data augmentation and a shallow ConvNet, and find that this baseline is surprisingly competitive with recent approaches (PVR, MVP, R3M…

2023

Visual Reinforcement Learning With Self-Supervised 3D Representations

RA-L 2023

A prominent approach to visual Reinforcement Learning (RL) is to learn an internal state representation using self-supervised methods, which has the potential benefit of improved sample-efficiency and generalization through additional learning signal and inductive biases. However, while the real wor

Cited by 74SourcecodeScholar
2022

UKPGAN: A General Self-Supervised Keypoint Detector

CVPR 2022poster

Keypoint detection is an essential component for the object registration and alignment. In this work, we reckon keypoint detection as information compression, and force the model to distill out important points of an object. Based on this, we propose UKPGAN, a general self-supervised 3D keypoint det…

Cited by 31PDFcodeScholar