← Search

Rujie Zhong

1 accepted papers

2022

Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement Learning

NeurIPS 2022accept

Reinforcement learning (RL) algorithms are often categorized as either on-policy or off-policy depending on whether they use data from a target policy of interest or from a different behavior policy. In this paper, we study a subtle distinction between on-policy data and on-policy sampling in the c…