Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning
Liangpeng Zhang, Ke Tang, Xin Yao
Abstract
Under/overestimation of state/action values are harmful for reinforcement learning agents. In this paper, we show that a state/action value estimated using the Bellman equation can be decomposed to a weighted sum of path-wise values that follow log-normal distributions. Since log-normal distributions are skewed, the distribution of estimated state/action values can also be skewed, leading to an imbalanced likelihood of under/overestimation. The degree of such imbalance can vary greatly among actions and policies within a single problem instance, making the agent prone to select actions/policies that have inferior expected return and higher likelihood of overestimation. We present a comprehensive analysis to such skewness, examine its factors and impacts through both theoretical and empirical results, and discuss the possible ways to reduce its undesirable effects.
BibTeX
@inproceedings{NIPS2017_69a5b599,
author = {Zhang, Liangpeng and Tang, Ke and Yao, Xin},
booktitle = {Advances in Neural Information Processing Systems},
editor = {I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning},
url = {https://proceedings.neurips.cc/paper_files/paper/2017/file/69a5b5995110b36a9a347898d97a610e-Paper.pdf},
volume = {30},
year = {2017}
}