2026
CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning
ICML 2026spotlight
Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as constrained Markov decision processes. While primal-dual methods scale well to deep RL, they often suffer from delayed constraint correction, leading to oscillatory behavi…