RA-L 20251 citations

Maximum Next-State Entropy for Efficient Reinforcement Learning

Dianyu Zhong, Yiqin Yang, Ziyou Zhang, Yuhua Jiang, Bo Xu, Qianchuan Zhao

Abstract

Entropy regularization is widely used to improve policy optimization and encourage exploration in reinforcement learning. By maximizing both the expected return and entropy, the agent aims to succeed at the task while acting as randomly as possible. However, current methods based on policy entropy encourage the agent to explore diverse actions, but they do not directly promote exploring diverse states. In this study, we theoretically reveal the challenge of optimizing the agent's next-state entropy and the gap between maximum next-state entropy and current policy entropy regularization methods. To address this limitation, we introduce <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">M</b>aximum <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">N</b>ext-<bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">S</b>tate <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E</b>ntropy (<bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">MNSE</b>), a novel method that maximizes next-state entropy through an action mapping layer following the inner policy. We provide a theoretical analysis demonstrating that MNSE can maximize next-state entropy by optimizing the entropy of the inner policy. We conduct extensive experiments on various continuous control tasks and demonstrate that MNSE can significantly improve the performance of RL algorithms.

BibTeX
@inproceedings{ral2025_maximumnextstate,
  title = {Maximum Next-State Entropy for Efficient Reinforcement Learning},
  author = {Dianyu Zhong and Yiqin Yang and Ziyou Zhang and Yuhua Jiang and Bo Xu and Qianchuan Zhao},
  booktitle = {RA-L 2025},
  year = {2025}
}
Maximum Next-State Entropy for Efficient Reinforcement Learning · RA-L 2025