← Search

Yaosheng Xu

1 accepted papers

2022

Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks

ICML 2022spotlight

In temporal-difference reinforcement learning algorithms, variance in value estimation can cause instability and overestimation of the maximal target value. Many algorithms have been proposed to reduce overestimation, including several recent ensemble methods, however none have shown success in samp…