NeurIPS 2020oral70 citations
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, Mengxiao Zhang
Abstract
We develop a new approach to obtaining high probability regret bounds for online learning with bandit feedback against an adaptive adversary. While existing approaches all require carefully constructing optimistic and biased loss estimators, our approach uses standard unbiased estimators and relies on a simple increasing learning rate schedule, together with the help of logarithmically homogeneous self-concordant barriers and a strengthened Freedman's inequality.
BibTeX
@inproceedings{NEURIPS2020_b2ea5e97,
author = {Lee, Chung-Wei and Luo, Haipeng and Wei, Chen-Yu and Zhang, Mengxiao},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
pages = {15522--15533},
publisher = {Curran Associates, Inc.},
title = {Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs},
url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/b2ea5e977c5fc1ccfa74171a9723dd61-Paper.pdf},
volume = {33},
year = {2020}
}