NeurIPS 2020poster90 citations

Minimax Value Interval for Off-Policy Evaluation and Policy Optimization

Nan Jiang, Jiawei Huang

Abstract

We study minimax methods for off-policy evaluation (OPE) using value functions and marginalized importance weights. Despite that they hold promises of overcoming the exponential variance in traditional importance sampling, several key problems remain: (1) They require function approximation and are generally biased. For the sake of trustworthy OPE, is there anyway to quantify the biases? (2) They are split into two styles (“weight-learning” vs “value-learning”). Can we unify them? In this paper we answer both questions positively. By slightly altering the derivation of previous methods (one from each style), we unify them into a single value interval that comes with a special type of double robustness: when either the value-function or the importance-weight class is well specified, the interval is valid and its length quantifies the misspecification of the other class. Our interval also provides a unified view of and new insights to some recent methods, and we further explore the implications of our results on exploration and exploitation in off-policy policy optimization with insufficient data coverage.

BibTeX
@inproceedings{NEURIPS2020_1cd138d0,
 author = {Jiang, Nan and Huang, Jiawei},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
 pages = {2747--2758},
 publisher = {Curran Associates, Inc.},
 title = {Minimax Value Interval for Off-Policy Evaluation and Policy Optimization},
 url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/1cd138d0499a68f4bb72bee04bbec2d7-Paper.pdf},
 volume = {33},
 year = {2020}
}