NeurIPS 2022accept20 citations

Conformal Off-Policy Prediction in Contextual Bandits

Muhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh, Arnaud Doucet

Abstract

Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, the expectation may not be the best measure of performance as it does not capture the variability of the outcome. In addition, particularly in safety-critical settings, stronger guarantees than asymptotic correctness may be required. To address these limitations, we consider a novel application of conformal prediction to contextual bandits. Given data collected under a behavioral policy, we propose \emph{conformal off-policy prediction} (COPP), which can output reliable predictive intervals for the outcome under a new target policy. We provide theoretical finite-sample guarantees without making any additional assumptions beyond the standard contextual bandit setup, and empirically demonstrate the utility of COPP compared with existing methods on synthetic and real-world data.

conformal predictioncontextual banditsuncertainty quantificationrobust ML
BibTeX
@inproceedings{
taufiq2022conformal,
title={Conformal Off-Policy Prediction in Contextual Bandits},
author={Muhammad Faaiz Taufiq and Jean-Francois Ton and Rob Cornish and Yee Whye Teh and Arnaud Doucet},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=IfgOWI5v2f}
}
Conformal Off-Policy Prediction in Contextual Bandits · NeurIPS 2022