2020
Infinite-horizon Off-Policy Policy Evaluation with Multiple Behavior Policies
ICLR 2020poster
We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite horizon context for off-policy policy evaluation. We propose estim…