NeurIPS 2025poster0 citations

CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement Learning

Ao Zhou, Jiayi Guan, Li Shen, Fan Lu, Sanqing Qu, Junqiao Zhao, Ziqiao Wang, Ya Wu

Abstract

Constrained hybrid-action reinforcement learning (RL) promises to learn a safe policy within a parameterized action space, which is particularly valuable for safety-critical applications involving discrete-continuous hybrid action spaces. However, existing hybrid-action RL algorithms primarily focus on reward maximization, which faces significant challenges for tasks involving both cost constraints and hybrid action spaces. In this work, we propose a novel Constrained Hybrid-action Policy Optimization algorithm (CHPO) to address the problems of constrained hybrid-action RL. Concretely, we rethink the limitations of hybrid-action RL in handling safe tasks with parameterized action spaces and reframe the objective of constrained hybrid-action RL by introducing the concept of Constrained Parameterized-action Markov Decision Process (CPMDP). Subsequently, we present a constrained hybrid-action policy optimization algorithm to confront the constrained hybrid-action problems and conduct theoretical analyses demonstrating that the CHPO converges to the optimal solution while satisfying safety constraints. Finally, extensive experiments demonstrate that the CHPO achieves competitive performance across multiple experimental tasks.

Reinforcement learningparameterized action spacesconstrained policy optimization.
BibTeX
@inproceedings{
zhou2025chpo,
title={{CHPO}: Constrained Hybrid-action Policy Optimization for Reinforcement Learning},
author={Ao Zhou and Jiayi Guan and Li Shen and Fan Lu and Sanqing Qu and Junqiao Zhao and Ziqiao Wang and Ya Wu and Guang Chen},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=KMaBmPHkVj}
}
CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement Learning · NeurIPS 2025