← Search

Abhishek Naik

2 accepted papers

2021

Learning and Planning in Average-Reward Markov Decision Processes

ICML 2021spotlight

We introduce learning and planning algorithms for average-reward MDPs, including 1) the first general proven-convergent off-policy model-free control algorithm without reference states, 2) the first proven-convergent off-policy model-free prediction algorithm, and 3) the first off-policy learning al…