AAAI 2023technical8 citations

State-Conditioned Adversarial Subgoal Generation

Vivienne Huiling Wang, Joni Pajarinen, Tinghuai Wang, Joni-Kristian Kämäräinen

Abstract

Hierarchical reinforcement learning (HRL) proposes to solve difficult tasks by performing decision-making and control at successively higher levels of temporal abstraction. However, off-policy HRL often suffers from the problem of a non-stationary high-level policy since the low-level policy is constantly changing. In this paper, we propose a novel HRL approach for mitigating the non-stationarity by adversarially enforcing the high-level policy to generate subgoals compatible with the current instantiation of the low-level policy. In practice, the adversarial learning is implemented by training a simple state conditioned discriminator network concurrently with the high-level policy which determines the compatibility level of subgoals. Comparison to state-of-the-art algorithms shows that our approach improves both learning efficiency and performance in challenging continuous control tasks.

BibTeX
@article{Wang_Pajarinen_Wang_Kämäräinen_2023, title={State-Conditioned Adversarial Subgoal Generation}, volume={37}, url={https://ojs.aaai.org/index.php/AAAI/article/view/26213}, DOI={10.1609/aaai.v37i8.26213}, abstractNote={Hierarchical reinforcement learning (HRL) proposes to solve difficult tasks by performing decision-making and control at successively higher levels of temporal abstraction. However, off-policy HRL often suffers from the problem of a non-stationary high-level policy since the low-level policy is constantly changing. In this paper, we propose a novel HRL approach for mitigating the non-stationarity by adversarially enforcing the high-level policy to generate subgoals compatible with the current instantiation of the low-level policy. In practice, the adversarial learning is implemented by training a simple state conditioned discriminator network concurrently with the high-level policy which determines the compatibility level of subgoals. Comparison to state-of-the-art algorithms shows that our approach improves both learning efficiency and performance in challenging continuous control tasks.}, number={8}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Wang, Vivienne Huiling and Pajarinen, Joni and Wang, Tinghuai and Kämäräinen, Joni-Kristian}, year={2023}, month={Jun.}, pages={10184-10191} }