← Search

Miroslav Dudı́k

1 accepted papers

2017

Optimal and Adaptive Off-policy Evaluation in Contextual Bandits

ICML 2017poster

We study the off-policy evaluation problem—estimating the value of a target policy using data collected by another policy—under the contextual bandit model. We consider the general (agnostic) setting without access to a consistent model of rewards and establish a minimax lower bound on the mean squa…

Cited by 254SourcePDFScholar