← Search

Kaizhe Hu

13 accepted papers

2026

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence

CVPR 2026

Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision foundation models (VFMs), the proposed AffordGen framework overcome

Cited by 0SourceScholar
2026

Failure-Aware RL: Reliable Offline-To-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

ICRA 2026poster

Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real…

2025

DenseMatcher: Learning 3D Semantic Correspondence for Category-Level Manipulation from a Single Demo

ICLR 2025spotlight

Dense 3D correspondence can enhance robotic manipulation by enabling the generalization of spatial, functional, and dynamic information from one object to an unseen counterpart. Compared to shape correspondence, semantic correspondence is more effective in generalizing across different object catego…

2025

Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

CoRL 2025poster

Simulation-based reinforcement learning (RL) has significantly advanced humanoid locomotion tasks, yet direct real-world RL from scratch or starting from pretrained policies remains rare, limiting the full potential of humanoid robots. Real-world training, despite being crucial for overcoming the si…

Cited by 0SourceScholar
2025

Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion Inversion

ICLR 2025spotlight

Visual imitation learning methods demonstrate strong performance, yet they lack generalization when faced with visual input perturbations like variations in lighting and textures. This limitation hampers their practical application in real-world settings. To address this, we propose ***Stem-OB*** th…

2024

Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion

NeurIPS 2024poster

Can we generate a control policy for an agent using just one demonstration of desired behaviors as a prompt, as effortlessly as creating an image from a textual description? In this paper, we present **Make-An-Agent**, a novel policy parameter generator that leverages the power of conditional diffus…

2024

Rethinking Transformers in Solving POMDPs

ICML 2024poster

Sequential decision-making algorithms such as reinforcement learning (RL) in real-world scenarios inevitably face environments with partial observability. This paper scrutinizes the effectiveness of a popular architecture, namely Transformers, in Partially Observable Markov Decision Processes (POMDP…

2024

Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation

ECCV 2024poster

"Enabling robotic manipulation that generalizes to out-of-distribution scenes is a crucial step toward the open-world embodied intelligence. For human beings, this ability is rooted in the understanding of semantic correspondence among different objects, which helps to naturally transfer the interac…

Cited by 48SourcePDFScholar
2024

Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization

ICLR 2024poster

Combining offline and online reinforcement learning (RL) is crucial for efficient and safe learning. However, previous approaches treat offline and online learning as separate procedures, resulting in redundant designs and limited performance. We ask: *Can we achieve straightforward yet effective of…

2023

Can Pre-Trained Text-to-Image Models Generate Visual Goals for Reinforcement Learning?

NeurIPS 2023poster

Pre-trained text-to-image generative models can produce diverse, semantically rich, and realistic images from natural language descriptions. Compared with language, images usually convey information with more details and less ambiguity. In this study, we propose Learning from the Void (LfVoid), a me…

2023

RL-ViGen: A Reinforcement Learning Benchmark for Visual Generalization

NeurIPS 2023poster

Visual Reinforcement Learning (Visual RL), coupled with high-dimensional observations, has consistently confronted the long-standing challenge of out-of-distribution generalization. Despite the focus on algorithms aimed at resolving visual generalization problems, we argue that the devil is in the e…