← Search

May Wang

1 accepted papers

2018

Variance Regularized Counterfactual Risk Minimization via Variational Divergence Minimization

ICML 2018oral

Off-policy learning, the task of evaluating and improving policies using historic data collected from a logging policy, is important because on-policy evaluation is usually expensive and has adverse impacts. One of the major challenge of off-policy learning is to derive counterfactual estimators tha…