NeurIPS 2019oral182 citations

Batched Multi-armed Bandits Problem

Zijun Gao, Yanjun Han, Zhimei Ren, Zhengqing Zhou

Abstract

In this paper, we study the multi-armed bandit problem in the batched setting where the employed policy must split data into a small number of batches. While the minimax regret for the two-armed stochastic bandits has been completely characterized in \cite{perchet2016batched}, the effect of the number of arms on the regret for the multi-armed case is still open. Moreover, the question whether adaptively chosen batch sizes will help to reduce the regret also remains underexplored. In this paper, we propose the BaSE (batched successive elimination) policy to achieve the rate-optimal regrets (within logarithmic factors) for batched multi-armed bandits, with matching lower bounds even if the batch sizes are determined in an adaptive manner.

BibTeX
@inproceedings{NEURIPS2019_20f07591,
 author = {Gao, Zijun and Han, Yanjun and Ren, Zhimei and Zhou, Zhengqing},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {Batched Multi-armed Bandits Problem},
 url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/20f07591c6fcb220ffe637cda29bb3f6-Paper.pdf},
 volume = {32},
 year = {2019}
}