Adaptive Motion Priors with Constrained Optimization
Tuchapong Sangthaworn, Bawornsak Sakulkueakulsuk
Abstract
Choosing locomotion learning paradigm in high-DOF system like humanoid robot faces several challenges. Free exploration creates complex reward surfaces that resist efficient exploration, while human motion priors cannot be directly copied due to different mechanical constraints. We present Adaptive Motion Priors with Constrained Optimization (AMPCO), a novel framework that transitions from human reference motions to task-focused optimization within learned behavioral bounds. AMPCO employs a two-phase optimization strategy: (1) Adaptive Imitation Guidance that prioritizes human motion, and (2) Adaptive Reward Weighting for Constrained Optimization that optimizes task objectives while maintaining motion quality within statistically-guaranteed bounds from Phase I. The transition between phases is automatically detected through percentile-based breakout detection from discriminator convergence. AMPCO introduces adaptive weighting mechanisms that smoothly adjust the importance of human imitation based on learning progress. Our experiments on the Unitree G1 humanoid robot simulation demonstrate that AMPCO reduces energy consumption variance by 67-90% across all baseline methods while achieving 70% lower energy consumption than task-focused baseline while maintaining velocity tracking accuracy comparable to the best-performing methods, with minimal computational overhead (<0.012% per training cycle).