IROS 20250 citations

Continuously Improved Reinforcement Learning for Automated Driving

Xuerun Yan, Zhexi Lian, Jia Hu, Yongwei Feng, Binyang Song, Haoran Wang

Abstract

Reinforcement Learning (RL) offers a promising solution to enable evolutionary automated driving. However, conventional RL methods often struggle with risk performance, as updated policies may fail to enhance performance or even lead to deterioration. To address this challenge, this research introduces a High Confidence Policy Improvement Reinforcement Learning-based (HCPI-RL) planner, designed to achieve the monotonic evolution of automated driving. The HCPI-RL planner features a novel RL policy update paradigm, ensuring that each newly learned policy outperforms previous policies, achieving monotonic performance enhancement. Hence, the proposed HCPI-RL planner has the following features: i) Evolutionary automated driving with guaranteed monotonic performance enhancement; ii) Capability of handling scenarios with emergency; iii) Enhanced decision-making optimality. Experimental results demonstrate that the proposed HCPI-RL planner enhances policy return by at least 20.1% and driving efficiency by at least 15.6%, compared to the conventional RL-based planners.

BibTeX
@inproceedings{iros2025_continuouslyimpr,
  title = {Continuously Improved Reinforcement Learning for Automated Driving},
  author = {Xuerun Yan and Zhexi Lian and Jia Hu and Yongwei Feng and Binyang Song and Haoran Wang},
  booktitle = {IROS 2025},
  year = {2025}
}