← Search

Xue Bin Peng

27 accepted papers

2025

CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control

ICLR 2025spotlight

Motion diffusion models and Reinforcement Learning (RL) based control for physics-based simulations have complementary strengths for human motion generation. The former is capable of generating a wide variety of motions, adhering to intuitive control such as text, while the latter offers physically…

2025

Generalizable Humanoid Manipulation with 3D Diffusion Policies

IROS 2025

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the difficulty of acquiring generalizable skills and the expensivenes

Cited by 40SourcecodeScholar
2025

Learning Smooth Humanoid Locomotion through Lipschitz-Constrained Policies

IROS 2025

Reinforcement learning combined with sim-to-real transfer offers a general framework for developing locomotion controllers for legged robots. To facilitate successful deployment in the real world, smoothing techniques, such as low-pass filters and smoothness rewards, are often employed to develop po

Cited by 49SourcecodeScholar
2025

TWIST: Teleoperated Whole-Body Imitation System

CoRL 2025poster

Teleoperating humanoid robots in a whole-body manner marks a fundamental step toward developing general-purpose robotic intelligence, with human motion providing an ideal interface for controlling all degrees of freedom. Yet, most current humanoid teleoperation systems fall short of enabling coordin…

Cited by 0SourceScholar
2024

DiffuseLoco: Real-Time Legged Locomotion Control with Diffusion from Offline Datasets

CoRL 2024poster

Offline learning at scale has led to breakthroughs in computer vision, natural language processing, and robotic manipulation domains. However, scaling up learning for legged robot locomotion, especially with multiple skills in a single policy, presents significant challenges for prior online reinfor…

Cited by 29SourceScholar
2024

Generating Human Interaction Motions in Scenes with Text Control

ECCV 2024poster

"We present , a text-controlled scene-aware motion generation method based on denoising diffusion models. Previous text-to-motion methods focus on characters in isolation without considering scenes due to the limited availability of datasets that include motion, text descriptions, and interactive sc…

Cited by 40SourcePDFScholar
2024

HiLMa-Res: A General Hierarchical Framework via Residual RL for Combining Quadrupedal Locomotion and Manipulation

IROS 2024poster

This work presents HiLMa-Res, a hierarchical framework leveraging reinforcement learning to tackle manipulation tasks while performing continuous locomotion using quadrupedal robots. Unlike most previous efforts that focus on solving a specific task, HiLMa-Res is designed to be general for various l…

Cited by 3SourceScholar
2023

Creating a Dynamic Quadrupedal Robotic Goalkeeper with Reinforcement Learning

IROS 2023poster

We present a reinforcement learning (RL) framework that enables quadrupedal robots to perform soccer goalkeeping tasks in the real world. Soccer goalkeeping with quadrupeds is a challenging problem, that combines highly dynamic locomotion with precise and fast non-prehensile object (ball) manipulati…

Cited by 50SourceScholar
2023

Learning and Adapting Agile Locomotion Skills by Transferring Experience

RSS 2023poster

Legged robots have enormous potential in their range of capabilities, from navigating unstructured terrains to high-speed running. However, these capabilities bring with them difficult control problems, and designing controllers for highly agile dynamic motions remains a substantial challenge for ro…

2023

RoboPianist: Dexterous Piano Playing with Deep Reinforcement Learning

CoRL 2023poster

Replicating human-like dexterity in robot hands represents one of the largest open problems in robotics. Reinforcement learning is a promising approach that has achieved impressive progress in the last few years; however, the class of problems it has typically addressed corresponds to a rather narro…

Cited by 47SourcecodeScholar
2023

Robust and Versatile Bipedal Jumping Control through Reinforcement Learning

RSS 2023poster

This work aims to push the limits of agility for bipedal robots by enabling a torque-controlled bipedal robot to perform robust and versatile dynamic jumps in the real world. We present a reinforcement learning framework for training a robot to accomplish a large variety of jumping tasks, such as ju…

Cited by 42SourcePDFScholar
2023

Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion

CVPR 2023poster

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability of trajectories, which is normally only associated with rule…

Cited by 118SourcePDFScholar
2023

Video Prediction Models as Rewards for Reinforcement Learning

NeurIPS 2023poster

Specifying reward signals that allow agents to learn complex behaviors is a long-standing challenge in reinforcement learning. A promising approach is to extract preferences for behaviors from unlabeled videos, which are widely available on the internet. We present Video Prediction Rewards (VIPER),…

Cited by 67SourcePDFScholar
2022

Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions

IROS 2022poster

Training a high-dimensional simulated agent with an under-specified reward function often leads the agent to learn physically infeasible strategies that are ineffective when deployed in the real world. To mitigate these unnatural behaviors, reinforcement learning practitioners often utilize complex…

Cited by 123SourceScholar
2022

GenLoco: Generalized Locomotion Controllers for Quadrupedal Robots

CoRL 2022poster

Recent years have seen a surge in commercially-available and affordable quadrupedal robots, with many of these platforms being actively used in research and industry. As the availability of legged robots grows, so does the need for controllers that enable these robots to perform useful skills. Howev…

Cited by 70SourcecodeScholar
2022

Hierarchical Reinforcement Learning for Precise Soccer Shooting Skills using a Quadrupedal Robot

IROS 2022poster

We address the problem of enabling quadrupedal robots to perform precise shooting skills in the real world using reinforcement learning. Developing algorithms to enable a legged robot to shoot a soccer ball to a given target is a challenging problem that combines robot motion control and planning in…

Cited by 68SourceScholar
2022

Legged Robots that Keep on Learning: Fine-Tuning Locomotion Policies in the Real World

ICRA 2022poster

Legged robots are physically capable of traversing a wide range of challenging environments, but designing controllers that are sufficiently robust to handle this diversity has been a long-standing challenge in robotics. Reinforcement learning presents an appealing approach for automating the contro…

Cited by 136SourceScholar
2022

Unsupervised Reinforcement Learning with Contrastive Intrinsic Control

NeurIPS 2022accept

We introduce Contrastive Intrinsic Control (CIC), an unsupervised reinforcement learning (RL) algorithm that maximizes the mutual information between state-transitions and latent skill vectors. CIC utilizes contrastive learning between state-transitions and skills vectors to learn behaviour embeddin…

Cited by 42SourcePDFScholar
2021

Offline Meta-Reinforcement Learning with Advantage Weighting

ICML 2021spotlight

This paper introduces the offline meta-reinforcement learning (offline meta-RL) problem setting and proposes an algorithm that performs well in this setting. Offline meta-RL is analogous to the widely successful supervised learning strategy of pre-training a model on a large batch of fixed, pre-coll…

2021

Reinforcement Learning for Robust Parameterized Locomotion Control of Bipedal Robots

ICRA 2021poster

Developing robust walking controllers for bipedal robots is a challenging endeavor. Traditional model-based locomotion controllers require simplifying assumptions and careful modelling; any small errors can result in unstable control. To address these challenges for bipedal locomotion, we present a…

Cited by 287SourceScholar
2020

Learning Agile Robotic Locomotion Skills by Imitating Animals

RSS 2020poster

Reproducing the diverse and agile locomotion skills of animals has been a longstanding challenge in robotics. While manually-designed controllers have been able to emulate many complex behaviors, building such controllers involves a time-consuming and difficult development process, often requiring s…

Cited by 614SourcePDFScholar
2020

Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives

ICLR 2020poster

Reinforcement learning agents that operate in diverse and complex environments can benefit from the structured decomposition of their behavior. Often, this is addressed in the context of hierarchical reinforcement learning, where the aim is to decompose a policy into lower-level primitives or option…

Cited by 55SourceScholar
2019

MCP: Learning Composable Hierarchical Control with Multiplicative Compositional Policies

NeurIPS 2019poster

Humans are able to perform a myriad of sophisticated tasks by drawing upon skills acquired through prior experience. For autonomous agents to have this capability, they must be able to extract reusable skills from past experience that can be recombined in new ways for subsequent tasks. Furthermore,…

2019

Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow

ICLR 2019poster

Adversarial learning methods have been proposed for a wide range of applications, but the training of adversarial models can be notoriously unstable. Effectively balancing the performance of the generator and discriminator is critical, since a discriminator that achieves very high accuracy will prod…

Cited by 282SourcePDFScholar
2018

Sim-to-Real Transfer of Robotic Control with Dynamics Randomization

ICRA 2018poster

Simulations are attractive environments for training agents as they provide an abundant source of data and alleviate certain safety concerns during the training process. But the behaviours developed by agents in simulation are often specific to the characteristics of the simulator. Due to modeling e…

Cited by 1799SourceScholar