2015
Off-policy Model-based Learning under Unknown Factored Dynamics
ICML 2015poster
Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority without testing the new policy? To answer this question, we introduce the G-SCOPE algorithm that evaluates a new policy based…