ICASSP 2022accepted0 citations

Efficient and Stable Information Directed Exploration for Continuous Reinforcement Learning

Mingzhe Chen, Xi Xiao, Wanpeng Zhang, Xiaotian Gao

Abstract

In this paper, we investigate the exploration-exploitation dilemma of reinforcement learning algorithms. We adapt the information directed sampling, an exploration framework that measures the information gain of a policy, to the continuous reinforcement learning. To stabilize the off-policy learning process and further improve the sample efficiency, we propose to use a randomized learning target and to dynamically adjust the update-to-data ratio for different parts of the neural network model. Experiments show that our approach significantly improves over existing methods and successfully completes tasks with highly sparse reward signals.

BibTeX
@inproceedings{icassp2022_efficientandstab,
  title = {Efficient and Stable Information Directed Exploration for Continuous Reinforcement Learning},
  author = {Mingzhe Chen and Xi Xiao and Wanpeng Zhang and Xiaotian Gao},
  booktitle = {ICASSP 2022},
  year = {2022}
}