← Search

Matthew Soh

2 accepted papers

2019

Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

NeurIPS 2019poster

Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning. However, in practice, commonly used off-policy approximate dynamic programming methods based on Q-learning and actor-critic methods are highly sensitive to the data distribution…

Cited by 1291SourcePDFScholar