2022
A Non-asymptotic Analysis of Non-parametric Temporal-Difference Learning
NeurIPS 2022accept
Temporal-difference learning is a popular algorithm for policy evaluation. In this paper, we study the convergence of the regularized non-parametric TD(0) algorithm, in both the independent and Markovian observation settings. In particular, when TD is performed in a universal reproducing kernel Hilb…