NeurIPS 2015poster21 citations
Policy Evaluation Using the Ω-Return
Philip S. Thomas, Scott Niekum, Georgios Theocharous, George Konidaris
Abstract
We propose the Ω-return as an alternative to the λ-return currently used by the TD(λ) family of algorithms. The benefit of the Ω-return is that it accounts for the correlation of different length returns. Because it is difficult to compute exactly, we suggest one way of approximating the Ω-return. We provide empirical studies that suggest that it is superior to the λ-return and γ-return for a variety of problems.
BibTeX
@inproceedings{NIPS2015_0e65972d,
author = {Thomas, Philip S. and Niekum, Scott and Theocharous, Georgios and Konidaris, George},
booktitle = {Advances in Neural Information Processing Systems},
editor = {C. Cortes and N. Lawrence and D. Lee and M. Sugiyama and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Policy Evaluation Using the \Omega -Return},
url = {https://proceedings.neurips.cc/paper_files/paper/2015/file/0e65972dce68dad4d52d063967f0a705-Paper.pdf},
volume = {28},
year = {2015}
}