Scalable Multi-Objective Robot Reinforcement Learning through Gradient Conflict Resolution
Reinforcement Learning (RL) robot controllers usually aggregate many task objectives into one scalar reward. While large-scale proximal policy optimisation (PPO) has enabled impressive results such as robust real-world robot locomotion, many tasks still require careful reward tuning and remain britt…