Weighted Linear Bandits for Non-Stationary Environments
Yoan Russac, Claire Vernade, Olivier Cappé
Abstract
We consider a stochastic linear bandit model in which the available actions correspond to arbitrary context vectors whose associated rewards follow a non-stationary linear regression model. In this setting, the unknown regression parameter is allowed to vary in time. To address this problem, we propose D-LinUCB, a novel optimistic algorithm based on discounted linear regression, where exponential weights are used to smoothly forget the past. This involves studying the deviations of the sequential weighted least-squares estimator under generic assumptions. As a by-product, we obtain novel deviation results that can be used beyond non-stationary environments. We provide theoretical guarantees on the behavior of D-LinUCB in both slowly-varying and abruptly-changing environments. We obtain an upper bound on the dynamic regret that is of order d B
BibTeX
@inproceedings{NEURIPS2019_263fc48a,
author = {Russac, Yoan and Vernade, Claire and Capp\'{e}, Olivier},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Weighted Linear Bandits for Non-Stationary Environments},
url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/263fc48aae39f219b4c71d9d4bb4aed2-Paper.pdf},
volume = {32},
year = {2019}
}