2024
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
ICLR 2024spotlight
We consider off-policy evaluation (OPE) of deterministic target policies for reinforcement learning (RL) in environments with continuous action spaces. While it is common to use importance sampling for OPE, it suffers from high variance when the behavior policy deviates significantly from the target…