2017
Optimal and Adaptive Off-policy Evaluation in Contextual Bandits
ICML 2017poster
We study the off-policy evaluation problem—estimating the value of a target policy using data collected by another policy—under the contextual bandit model. We consider the general (agnostic) setting without access to a consistent model of rewards and establish a minimax lower bound on the mean squa…