ICASSP 2020accepted0 citations

Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions

Liangyu Huo, Zulin Wang, Mai Xu, Yuhang Song

Abstract

Automatic sub-policy discovery has recently received much attention in hierarchical reinforcement learning (HRL). The conventional approaches to learning sub-policies suffer from collapsing into just one sub-policy dominating the whole task, lacking techniques to ensure the diversity of different subpolicies. In this paper, we formulate the discovery of diverse sub-policies as a trajectory inference. Then, we propose an information-theoretic objective based on action distributions to encourage diversity. Moreover, two simplifications are derived on discrete and continuous action space for reducing the computation. Finally, the experimental results show that the proposed approach can further improve the state-of-theart approaches without modifying existing hyperparameters on two different HRL domains, suggesting the wide applicability and robustness of our approach.

BibTeX
@inproceedings{icassp2020_learningdiverses,
  title = {Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions},
  author = {Liangyu Huo and Zulin Wang and Mai Xu and Yuhang Song},
  booktitle = {ICASSP 2020},
  year = {2020}
}
Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions · ICASSP 2020