2026
Persistent Safety Set Guided Offline Safe Reinforcement Learning
IJCAI 2026
Offline safe reinforcement learning learns high-return policies that satisfy hard safety constraints using only a pre-collected dataset. This setting is challenging due to the inability to explore, and the risk of propagating value errors through unsafe state-space regions. To address this, first, w