← Search

OpenAI Jonathan Ho

2 accepted papers

2018

Evolved Policy Gradients

NeurIPS 2018spotlight

We propose a metalearning approach for learning gradient-based reinforcement learning (RL) algorithms. The idea is to evolve a differentiable loss function, such that an agent, which optimizes its policy to minimize this loss, will achieve high rewards. The loss is parametrized via temporal convolut…

2017

One-Shot Imitation Learning

NeurIPS 2017poster

Imitation learning has been commonly applied to solve different tasks in isolation. This usually requires either careful feature engineering, or a significant number of samples. This is far from what we desire: ideally, robots should be able to learn from very few demonstrations of any given task, a…

Cited by 870SourcePDFScholar