← Search

Zhenyu Jiang

17 accepted papers

2026

A Mechanistic Understanding of Sim-and-Real Co-Training in Generative Policies

ICML 2026poster

Co-training, which combines limited in-domain real-world data with abundant surrogate data such as simulation or cross-embodiment demonstrations, has been widely adopted for training generative visuomotor robot policies. Despite its empirical success, the mechanisms underlying when and why co-traini…

Cited by 0SourceScholar
2026

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos

ICRA 2026poster

We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples. In-context learning (ICL) is a promising framework for achieving this goal due to its test-time data efficiency and rapid adaptability. However, current ICL methods rely on labor-intensive teleope…

2026

Residual Off-Policy RL for Finetuning Behavior Cloning Policies

ICRA 2026poster

Recent advances in behavior cloning (BC) have enabled impressive visuomotor control policies. However, these approaches are limited by the quality of human demonstrations, the manual effort required for data collection, and the diminishing returns from offline data. In comparison, reinforcement lear…

2025

DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning

ICRA 2025

Imitation learning from human demonstrations is an effective means to teach robots manipulation skills. But data acquisition is a major bottleneck in applying this paradigm more broadly, due to the high costs and human efforts involved. There has been significant interest in imitation learning for b

Cited by 121SourcecodeScholar
2025

HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

ICRA 2025

Humanoid whole-body control requires adapting to diverse tasks such as navigation, loco-manipulation, and tabletop manipulation, each demanding a different mode of control. For example, navigation relies on root velocity or position tracking, while tabletop manipulation prioritizes upper-body joint

Cited by 126SourceScholar
2025

Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation

RSS 2025poster

Large real-world robot datasets hold great potential for developing generalist robot policies, but scaling real-world data collection is time-consuming, costly, and resource-intensive. Simulation offers a promising solution, with recent advances in generative AI and synthetic data generation tools e…

Cited by 4PDFScholar
2024

ARDuP: Active Region Video Diffusion for Universal Policies

IROS 2024poster

Sequential decision-making can be formulated as a text-conditioned video generation problem, where a video planner, guided by a text-defined goal, generates future frames visualizing planned actions, from which control actions are subsequently derived. In this work, we introduce Active Region Video…

Cited by 3SourceScholar
2024

Doduo: Learning Dense Visual Correspondence from Unsupervised Semantic-Aware Flow

ICRA 2024poster

Dense visual correspondence plays a vital role in robotic perception. This work focuses on establishing the dense correspondence between a pair of images that captures dynamic scenes undergoing substantial transformations. We introduce Doduo to learn general dense visual correspondence from in-the-w…

Cited by 6SourcecodeScholar
2024

Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

CoRL 2024poster

Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications and exhibit human-like behaviors. This work focuses on generat…

Cited by 9SourcecodeScholar
2024

LEAP: Liberate Sparse-View 3D Modeling from Camera Poses

ICLR 2024poster

Are camera poses necessary for multi-view 3D modeling? Existing approaches predominantly assume access to accurate camera poses. While this assumption might hold for dense views, accurately estimating camera poses for sparse views is often elusive. Our analysis reveals that noisy estimated poses lea…

2024

OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

CoRL 2024poster

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting,…

Cited by 33SourcecodeScholar
2023

Ditto in the House: Building Articulation Models of Indoor Scenes through Interactive Perception

ICRA 2023poster

Virtualizing the physical world into virtual models has been a critical technique for robot navigation and planning in the real world. To foster manipulation with articulated objects in everyday life, this work explores building articulation models of indoor scenes through a robot's purposeful inter…

Cited by 35SourcecodeScholar
2023

Learning Generalizable Manipulation Policies with Object-Centric 3D Representations

CoRL 2023poster

We introduce GROOT, an imitation learning method for learning robust policies with object-centric and 3D priors. GROOT builds policies that generalize beyond their initial training conditions for vision-based manipulation. It constructs object-centric 3D representations that are robust toward backgr…

Cited by 48SourcecodeScholar
2022

ACID: Action-Conditional Implicit Visual Dynamics for Deformable Object Manipulation

RSS 2022poster

Manipulating volumetric deformable objects in the real world, like plush toys and pizza dough, bring substantial challenges due to infinite shape variations, non-rigid motions, and partial observability. We introduce ACID, an action-conditional visual dynamics model for volumetric deformable objects…

Cited by 41SourcePDFScholar
2021

Synergies Between Affordance and Geometry: 6-DoF Grasp Detection via Implicit Representations

RSS 2021poster

Grasp detection in clutter requires the robot to reason about the 3D scene from incomplete and noisy perception. In this work; we draw insight that 3D reconstruction and grasp learning are two intimately connected tasks; both of which require a fine-grained understanding of local geometry details. W…

Cited by 169SourcePDFScholar
2020

Deep Face Super-Resolution With Iterative Collaboration Between Attentive Recovery and Landmark Estimation

CVPR 2020poster

Recent works based on deep learning and facial priors have succeeded in super-resolving severely degraded facial images. However, the prior knowledge is not fully exploited in existing methods, since facial priors such as landmark and component maps are always estimated by low-resolution or coarsely…

Cited by 220PDFcodeScholar