2023
Semiparametrically Efficient Off-Policy Evaluation in Linear Markov Decision Processes
ICML 2023poster
We study semiparametrically efficient estimation in off-policy evaluation (OPE) where the underlying Markov decision process (MDP) is linear with a known feature map. We characterize the variance lower bound for regular estimators in the linear MDP setting and propose an efficient estimator whose va…