NeurIPS 2019poster13 citations
Thompson Sampling with Information Relaxation Penalties
Seungki Min, Costis Maglaras, Ciamac C. Moallemi
Abstract
We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoints. Analogous to TS, which, at each decision epoch pulls an arm that is best with respect to the randomly sampled parameters, our algorithms sample entire future reward realizations and take the corresponding best action. However, this is done in the presence of “penalties” that seek to compensate for the availability of future information.
BibTeX
@inproceedings{NEURIPS2019_e5b294b7,
author = {Min, Seungki and Maglaras, Costis and Moallemi, Ciamac C},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Thompson Sampling with Information Relaxation Penalties},
url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/e5b294b70c9647dcf804d7baa1903918-Paper.pdf},
volume = {32},
year = {2019}
}