← Search

Paul Trichelair

1 accepted papers

2019

Safe Policy Improvement with Baseline Bootstrapping

ICML 2019oral

This paper considers Safe Policy Improvement (SPI) in Batch Reinforcement Learning (Batch RL): from a fixed dataset and without direct access to the true environment, train a policy that is guaranteed to perform at least as well as the baseline policy used to collect the data. Our approach, called…