2020
A new convergent variant of Q-learning with linear function approximation
NeurIPS 2020poster
In this work, we identify a novel set of conditions that ensure convergence with probability 1 of Q-learning with linear function approximation, by proposing a two time-scale variation thereof. In the faster time scale, the algorithm features an update similar to that of DQN, where the impact of boo…