2018
An Off-policy Policy Gradient Theorem Using Emphatic Weightings
NeurIPS 2018poster
Policy gradient methods are widely used for control in reinforcement learning, particularly for the continuous action setting. There have been a host of theoretically sound algorithms proposed for the on-policy setting, due to the existence of the policy gradient theorem which provides a simplified…