AISTATS 2020poster15 citations
Conditional Importance Sampling for Off-Policy Learning
Mark Rowland, Anna Harutyunyan, Hado Hasselt, Diana Borsa, Tom Schaul, Remi Munos, Will Dabney
Abstract
The principal contribution of this paper is a conceptual framework for off-policy reinforcement learning, based on conditional expectations of importance sampling ratios. This framework yields new perspectives and understanding of existing off-policy algorithms, and reveals a broad space of unexplored algorithms. We theoretically analyse this space, and concretely investigate several algorithms that arise from this framework.
BibTeX
@InProceedings{pmlr-v108-rowland20b,
title = {Conditional Importance Sampling for Off-Policy Learning},
author = {Rowland, Mark and Harutyunyan, Anna and van Hasselt, Hado and Borsa, Diana and Schaul, Tom and Munos, Remi and Dabney, Will},
booktitle = {Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics},
pages = {45--55},
year = {2020},
editor = {Chiappa, Silvia and Calandra, Roberto},
volume = {108},
series = {Proceedings of Machine Learning Research},
month = {26--28 Aug},
publisher = {PMLR},
pdf = {http://proceedings.mlr.press/v108/rowland20b/rowland20b.pdf},
url = {https://proceedings.mlr.press/v108/rowland20b.html},
abstract = {The principal contribution of this paper is a conceptual framework for off-policy reinforcement learning, based on conditional expectations of importance sampling ratios. This framework yields new perspectives and understanding of existing off-policy algorithms, and reveals a broad space of unexplored algorithms. We theoretically analyse this space, and concretely investigate several algorithms that arise from this framework.}
}