2026
SRPO: Self-Reflection Policy Optimization for Stable and Robust Autonomous Driving
ICRA 2026poster
Autonomous driving demands reinforcement learning (RL) agents that are not only performant, but also stable, sample-efficient, and robust to uncertainty. However, conventional policy optimization methods often suffer from unstable convergence, sensitivity to reward scaling, and limited generalizatio…