← Search

Shunpeng Yang

7 accepted papers

2026

PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement Learning

ICLR 2026poster

Among on-policy reinforcement learning algorithms, Proximal Policy Optimization (PPO) demonstrates is widely favored for its simplicity, numerical stability, and strong empirical performance. Standard PPO relies on surrogate objectives defined via importance ratios, which require evaluating policy l…

Cited by 0SourceScholar
2025

Multi-Loco: Unifying Multi-Embodiment Legged Locomotion via Reinforcement Learning Augmented Diffusion

CoRL 2025poster

Generalizing locomotion policies across diverse legged robots with varying morphologies is a key challenge due to differences in observation/action dimensions and system dynamics. In this work, we propose \textit{Multi-Loco}, a novel unified framework combining a morphology-agnostic generative diffu…

Cited by 0SourceScholar
2024

Task-Space Riccati Feedback based Whole Body Control for Underactuated Legged Locomotion

IROS 2024poster

This manuscript primarily aims to enhance the performance of whole-body controllers(WBC) for underactuated legged locomotion. We introduce a systematic parameter design mechanism for the floating-base feedback control within the WBC. The proposed approach involves utilizing the linearized model of u…

Cited by 0SourceScholar
2023

Template Model Inspired Task Space Learning for Robust Bipedal Locomotion

IROS 2023poster

This work presents a hierarchical framework for bipedal locomotion that combines a Reinforcement Learning (RL)-based high-level (HL) planner policy for the online generation of task space commands with a model-based low-level (LL) controller to track the desired task space trajectories. Different fr…

Cited by 16SourceScholar
2022

Improved Task Space Locomotion Controller for a Quadruped Robot with Parallel Mechanisms

IROS 2022poster

In this work, an advanced quadruped robot with abundant kinematic loops and passive joints is introduced. Due to the existence of many closed chains, the robot dynamic model is quite complex, and is derived using the Gauss's principle of least constraint. To explicitly consider the loop-closure cons…

Cited by 3SourceScholar
2021

Force-feedback based Whole-body Stabilizer for Position-Controlled Humanoid Robots

IROS 2021poster

This paper studies stabilizer design for position-controlled humanoid robots. Stabilizers are an essential part for position-controlled humanoids, whose primary objective is to adjust the control input sent to the robot to assist the tracking controller to better follow the planned reference traject…

Cited by 9SourceScholar
2021

Reachability-based Push Recovery for Humanoid Robots with Variable-Height Inverted Pendulum

ICRA 2021poster

This paper studies push recovery for humanoid robots based on a variable-height inverted pendulum (VHIP) model. We first develop an approach for treating zero-step capturability of the VHIP with a novel methodology based on Hamilton-Jacobi (HJ) reachability analysis. Such an approach uses the sub-ze…

Cited by 7SourceScholar