2026
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
ICML 2026poster
Safety is an essential requirement for reinforcement learning systems. The newly emerging framework of robust constrained Markov decision processes allows learning policies that satisfy long-term constraints while providing guarantees under epistemic uncertainty. This paper presents mirror descent p…