← Search

Mingqi Yuan

6 accepted papers

2026

Gait-Adaptive Perceptive Humanoid Locomotion With Real-Time Under-Base Terrain Reconstruction

RA-L 2026

For full-size humanoid robots, reliable locomotion on complex terrains—such as long staircases—remains challenging, even with recent advances in reinforcement-learning-based control. In such settings, limited perception, ambiguous terrain cues, and insufficient adaptation of gait timing can cause ev

Cited by 9SourcecodeScholar
2026

Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning

CVPR 2026

Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Designing such reward functions can be challenging and may not generalize well across different tasks. To address this limi

Cited by 0SourceScholar
2026

PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations

CVPR 2026

Achieving efficient and robust whole-body control (WBC) is essential for enabling humanoid robots to perform complex tasks in dynamic environments. Despite the success of reinforcement learning (RL) in this domain, its sample inefficiency remains a significant challenge due to the intricate dynamics

Cited by 0SourcecodeScholar
2025

RLLTE: Long-Term Evolution Project of Reinforcement Learning

AAAI 2025technical

We present RLLTE: a long-term evolution, extremely modular, and open-source framework for reinforcement learning (RL) research and application. Beyond delivering top-notch algorithm implementations, RLLTE also serves as a toolkit for developing algorithms. More specifically, RLLTE decouples the RL a…

2025

ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning

ICCV 2025poster

Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity a…

2023

Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement Learning

ICML 2023poster

We present AIRS: **A**utomatic **I**ntrinsic **R**eward **S**haping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task retu…