RA-L 20260 citations

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

Khang Nguyen, An T. Le, Jan Peters, Minh Nhat Vu

Abstract

Achieving robust robot learning for humanoid locomotion is a fundamental challenge in model-based reinforcement learning (MBRL), where environmental stochasticity and randomness can hinder efficient exploration and learning stability. The environmental, so-called aleatoric, uncertainty can be amplified in high-dimensional action spaces with complex contact dynamics and entangled with epistemic uncertainty in the models during learning phases. In this work, we propose <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DoublyAware</i>, an uncertainty-aware extension of Temporal Difference Model Predictive Control (TD-MPC) that explicitly decomposes uncertainty into two disjoint, interpretable components, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e.</i>, planning and policy uncertainties. To handle the planning uncertainty, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DoublyAware</i> employs conformal prediction to filter candidate trajectories using quantile-calibrated risk bounds, ensuring statistical consistency and robustness against stochastic dynamics. Meanwhile, policy rollouts are leveraged as structured informative priors to support the learning phase with Group-Relative Policy Constraint (GRPC) optimizers, which impose a group-based adaptive trust region in the latent action space. This combination enables the robot agent to prioritize high-confidence, high-reward behavior while maintaining effective, targeted exploration under uncertainty. Evaluated on the HumanoidBench locomotion suite with the Unitree 26-DoF H1-2 humanoid, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DoublyAware</i> demonstrates improved sample efficiency, accelerated convergence, and enhanced motion feasibility compared to RL baselines. Our results emphasize the significance of structured uncertainty modeling for data-efficient and reliable decision-making in TD-MPC-based humanoid locomotion learning.

BibTeX
@inproceedings{ral2026_doublyawaredualp,
  title = {DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion},
  author = {Khang Nguyen and An T. Le and Jan Peters and Minh Nhat Vu},
  booktitle = {RA-L 2026},
  year = {2026}
}
DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion · RA-L 2026