← Search

Chris Nota

2 accepted papers

2021

Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods

ICML 2021spotlight

Hindsight allows reinforcement learning agents to leverage new observations to make inferences about earlier states and transitions. In this paper, we exploit the idea of hindsight and introduce posterior value functions. Posterior value functions are computed by inferring the posterior distribution…

Cited by 12SourcePDFScholar