← Search

Yori Zwols

2 accepted papers

2019

Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search

ICLR 2019poster

Learning policies on data synthesized by models can in principle quench the thirst of reinforcement learning algorithms for large amounts of real experience, which is often costly to acquire. However, simulating plausible experience de novo is a hard problem for many complex environments, often resu…

Cited by 166SourcePDFScholar
2018

Generative Temporal Models with Spatial Memory for Partially Observed Environments

ICML 2018oral

In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent’s representations during training or via use as part of an explicit planning mechanism. However, their application in practice has been limite…

Cited by 32SourcePDFScholar