← Search

Liangpeng Zhang

2 accepted papers

2017

Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning

NeurIPS 2017poster

Under/overestimation of state/action values are harmful for reinforcement learning agents. In this paper, we show that a state/action value estimated using the Bellman equation can be decomposed to a weighted sum of path-wise values that follow log-normal distributions. Since log-normal distribution…

Cited by 6SourcePDFScholar