2015
Modelling Policies in MDPs in Reproducing Kernel Hilbert Space
AISTATS 2015poster
We consider modelling policies for MDPs in (vector-valued) reproducing kernel Hilbert function spaces (RKHS). This enables us to work “non-parametrically” in a rich function class, and provides the ability to learn complex policies. We present a framework for performing gradient-based policy optimiz…