← Search

Hugo Penedones

1 accepted papers

2019

Adaptive Temporal-Difference Learning for Policy Evaluation with Per-State Uncertainty Estimates

NeurIPS 2019poster

We consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (TD) learning and Monte Carlo (MC) policy evaluation. The two methods are known to achieve complementary bias-variance tra…

Cited by 10SourcePDFScholar