2022
Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks
ICML 2022spotlight
In temporal-difference reinforcement learning algorithms, variance in value estimation can cause instability and overestimation of the maximal target value. Many algorithms have been proposed to reduce overestimation, including several recent ensemble methods, however none have shown success in samp…