ICRA 2026poster0 citations

Teaching to Individual Needs: Bidirectional Teacher-Student Learning for Wheeled-Legged Locomotion

Guangsheng Li, Charles Wu, XinHua Zheng, Shiyu Zhu, Shenglan Liu

Abstract

Reinforcement Learning (RL) enables robust and adaptive locomotion in legged and wheeled-legged robots. A common approach is the Teacher-Student (TS) paradigm, in which a teacher policy with privileged information supervises a proprioceptive student. While the TS paradigm has proven effective on legged robots, we encounter two critical issues when applying it to wheeled-legged robots. One issue is multimodal confusion, where teacher actions become multimodal under the student proprioceptive observations, resulting in the student generating averaged action modes. The other is low imitability of teacher actions, as the teacher overlooks their reproducibility by the student. To address these issues, we propose Teaching to Individual Needs (TIN), a bidirectional TS framework. To mitigate multimodal confusion within the student policy, we design a Highest-Weight Component Mixture Density Network (HWC-MDN). By utilizing HWC-MDN, TIN student can explicitly model multimodal action distributions and outputs the highest-weight component. To improve imitability, we propose an Imitation-Aware Reward (IAR) that encourages the teacher to generate more reproducible actions by the student. Simulation experiments show that TIN significantly improves both training efficiency and traversability. Real-world tests illustrate that TIN enables the wheeled-legged robot MagicDog-W to traverse 45 cm obstacles and ascend 45° slopes.

Legged RobotsReinforcement LearningImitation Learning
Teaching to Individual Needs: Bidirectional Teacher-Student Learning for Wheeled-Legged Locomotion · ICRA 2026