A Maximum-Entropy Approach to Off-Policy Evaluation in Average-Reward MDPs
Nevena Lazic, Dong Yin, Mehrdad Farajtabar, Nir Levine, Dilan Gorur, Chris Harris, Dale Schuurmans
Abstract
This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs). For MDPs that are ergodic and linear (i.e. where rewards and dynamics are linear in some known features), we provide the first finite-sample OPE error bound, extending the existing results beyond the episodic and discounted cases. In a more general setting, when the feature dynamics are approximately linear and for arbitrary rewards, we propose a new approach for estimating stationary distributions with function approximation. We formulate this problem as finding the maximum-entropy distribution subject to matching feature expectations under empirical dynamics. We show that this results in an exponential-family distribution whose sufficient statistics are the features, paralleling maximum-entropy approaches in supervised learning. We demonstrate the effectiveness of the proposed OPE approaches in multiple environments.
BibTeX
@inproceedings{NEURIPS2020_9308b0d6,
author = {Lazic, Nevena and Yin, Dong and Farajtabar, Mehrdad and Levine, Nir and Gorur, Dilan and Harris, Chris and Schuurmans, Dale},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
pages = {12461--12471},
publisher = {Curran Associates, Inc.},
title = {A Maximum-Entropy Approach to Off-Policy Evaluation in Average-Reward MDPs},
url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/9308b0d6e5898366a4a986bc33f3d3e7-Paper.pdf},
volume = {33},
year = {2020}
}