2019
Adaptive Temporal-Difference Learning for Policy Evaluation with Per-State Uncertainty Estimates
NeurIPS 2019poster
We consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (TD) learning and Monte Carlo (MC) policy evaluation. The two methods are known to achieve complementary bias-variance tra…