← Search

Joao Carvalho

2 accepted papers

2023

Diminishing Return of Value Expansion Methods in Model-Based Reinforcement Learning

ICLR 2023poster

Model-based reinforcement learning is one approach to increase sample efficiency. However, the accuracy of the dynamics model and the resulting compounding error over modelled trajectories are commonly regarded as key limitations. A natural question to ask is: How much more sample efficiency can be…

2020

A Nonparametric Off-Policy Policy Gradient

AISTATS 2020poster

Reinforcement learning (RL) algorithms still suffer from high sample complexity despite outstanding recent successes. The need for intensive interactions with the environment is especially observed in many widely popular policy gradient algorithms that perform updates using on-policy samples. The pr…