NeurIPS 2020oral70 citations

Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs

Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, Mengxiao Zhang

Abstract

We develop a new approach to obtaining high probability regret bounds for online learning with bandit feedback against an adaptive adversary. While existing approaches all require carefully constructing optimistic and biased loss estimators, our approach uses standard unbiased estimators and relies on a simple increasing learning rate schedule, together with the help of logarithmically homogeneous self-concordant barriers and a strengthened Freedman's inequality.

BibTeX
@inproceedings{NEURIPS2020_b2ea5e97,
 author = {Lee, Chung-Wei and Luo, Haipeng and Wei, Chen-Yu and Zhang, Mengxiao},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
 pages = {15522--15533},
 publisher = {Curran Associates, Inc.},
 title = {Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs},
 url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/b2ea5e977c5fc1ccfa74171a9723dd61-Paper.pdf},
 volume = {33},
 year = {2020}
}
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs · NeurIPS 2020