2018
Variance Regularized Counterfactual Risk Minimization via Variational Divergence Minimization
ICML 2018oral
Off-policy learning, the task of evaluating and improving policies using historic data collected from a logging policy, is important because on-policy evaluation is usually expensive and has adverse impacts. One of the major challenge of off-policy learning is to derive counterfactual estimators tha…