← Search

Zhengrong Xue

12 accepted papers

2026

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence

CVPR 2026

Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision foundation models (VFMs), the proposed AffordGen framework overcome

Cited by 0SourceScholar
2026

H$^3$DP: Triply‑Hierarchical Diffusion Policy for Visuomotor Learning

ICLR 2026poster

Visuomotor policy learning has witnessed substantial progress in robotic manipulation, with recent approaches predominantly relying on generative models to model the action distribution. However, these methods often overlook the critical coupling between visual perception and action prediction. In t…

Cited by 0SourcecodeScholar
2026

MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery

ICRA 2026poster

Diffusion policies have emerged as a powerful framework for robotic visuomotor control, yet they often lack the robustness to recover from subtask failures in long-horizon, multi-stage tasks and their learned representations of observations are often difficult to interpret. In this work, we propose …

2025

DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning

RSS 2025poster

Visuomotor policies have shown great promise in robotic manipulation but often require substantial amounts of human-collected data for effective performance. A key reason underlying the data demands is their limited spatial generalization capability, which necessitates extensive data collection acro…

Cited by 8PDFScholar
2025

DemoSpeedup: Accelerating Visuomotor Policies via Entropy-Guided Demonstration Acceleration

CoRL 2025oral

Imitation learning has shown great promise in robotic manipulation, but the policy’s execution is often unsatisfactorily slow due to commonly tardy demonstrations collected by human operators. In this work, we present DemoSpeedup, a self-supervised method to accelerate visuomotor policy execution vi…

Cited by 0SourceScholar
2025

MENTOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning

ICML 2025poster

Visual deep reinforcement learning (RL) enables robots to acquire skills from visual input for unstructured tasks. However, current algorithms suffer from low sample efficiency, limiting their practical applicability. In this work, we present MENTOR, a method that improves both the *architecture* an…

2024

ArrayBot: Reinforcement Learning for Generalizable Distributed Manipulation through Touch

ICRA 2024poster

We present ArrayBot, a distributed manipulation system consisting of a 16 × 16 array of vertically sliding pillars integrated with tactile sensors. Functionally, ArrayBot is designed to simultaneously support, perceive, and manipulate the tabletop objects. Towards generalizable distributed manipulat…

Cited by 14SourceScholar
2024

Key-Grid: Unsupervised 3D Keypoints Detection using Grid Heatmap Features

NeurIPS 2024poster

Detecting 3D keypoints with semantic consistency is widely used in many scenarios such as pose estimation, shape registration and robotics. Currently, most unsupervised 3D keypoint detection methods focus on the rigid-body objects. However, when faced with deformable objects, the keypoints they iden…

Cited by 1SourcePDFScholar
2024

RiEMann: Near Real-Time SE(3)-Equivariant Robot Manipulation without Point Cloud Segmentation

CoRL 2024poster

We present RiEMann, an end-to-end near Real-time SE(3)-Equivariant Robot Manipulation imitation learning framework from scene point cloud input. Compared to previous methods that rely on descriptor field matching, RiEMann directly predicts the target actions for manipulation without any object segme…

Cited by 16SourceScholar
2023

USEEK: Unsupervised SE(3)-Equivariant 3D Keypoints for Generalizable Manipulation

ICRA 2023poster

Can a robot manipulate intra-category unseen objects in arbitrary poses with the help of a mere demonstration of grasping pose on a single object instance? In this paper, we try to address this intriguing challenge by using USEEK, an unsupervised SE(3)-equivariant keypoints method that enjoys alignm…

Cited by 29SourceScholar
2022

Pre-Trained Image Encoder for Generalizable Visual Reinforcement Learning

NeurIPS 2022accept

Learning generalizable policies that can adapt to unseen environments remains challenging in visual Reinforcement Learning (RL). Existing approaches try to acquire a robust representation via diversifying the appearances of in-domain observations for better generalization. Limited by the specific ob…

Cited by 85SourcePDFScholar