NeurIPS 2021spotlight44 citations

Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning

Alberto Maria Metelli, Alessio Russo, Marcello Restelli

Abstract

Importance Sampling (IS) is a widely used building block for a large variety of off-policy estimation and learning algorithms. However, empirical and theoretical studies have progressively shown that vanilla IS leads to poor estimations whenever the behavioral and target policies are too dissimilar. In this paper, we analyze the theoretical properties of the IS estimator by deriving a novel anticoncentration bound that formalizes the intuition behind its undesired behavior. Then, we propose a new class of IS transformations, based on the notion of power mean. To the best of our knowledge, the resulting estimator is the first to achieve, under certain conditions, two key properties: (i) it displays a subgaussian concentration rate; (ii) it preserves the differentiability in the target distribution. Finally, we provide numerical simulations on both synthetic examples and contextual bandits, in comparison with off-policy evaluation and learning baselines.

Importance SamplingOff-Policy EvaluationOff-Policy LearningSubgaussian
BibTeX
@inproceedings{
metelli2021subgaussian,
title={Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning},
author={Alberto Maria Metelli and Alessio Russo and Marcello Restelli},
booktitle={Advances in Neural Information Processing Systems},
editor={A. Beygelzimer and Y. Dauphin and P. Liang and J. Wortman Vaughan},
year={2021},
url={https://openreview.net/forum?id=_8vCV7AxPZ}
}
Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning · NeurIPS 2021