ICRA 2026poster0 citations

Beyond Domain Randomization: Safety Certificates for Reinforcement Learning

Paula Stocco, Francesco Micheli, Niklas Schmid, John Lygeros, Efe Balta

Abstract

With the growing acceptance of robotics in daily life there is a growing need for certifiably safe control policies. While simulation provides a safe training environment, policies often fail in sim-to-real transfer. We propose a data-driven certification framework for reinforcement learning based on Pick-to-Learn (P2L), a meta-algorithm that uses data preference ordering to compute probabilistic bounds on the satisfaction of application dependent properties of interest. Our results demonstrate that using P2L maintains high performance while distinguishing between policies that appear similar under domain randomization alone. This work offers a practical method for preparing safe reinforcement learning policies by providing formal safety guarantees prior to hardware deployment.

Robot SafetyPlanning under UncertaintyReinforcement Learning