2020
Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation
ICLR 2020spotlight
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estim…