IJCAI 2022poster5 citations

Dynamic Bandits with Temporal Structure

Qinyi Chen

Abstract

In this work, we study a dynamic multi-armed bandit (MAB) problem, where the expected reward of each arm evolves over time following an auto-regressive model. We present an algorithm whose per-round regret upper bound almost matches the regret lower bound, and numerically demonstrate its efficacy in adapting to the changing environment.

Machine Learning (ML): General
BibTeX
@inproceedings{ijcai2022p823,
  title     = {Dynamic Bandits with Temporal Structure},
  author    = {Chen, Qinyi},
  booktitle = {Proceedings of the Thirty-First International Joint Conference on
               Artificial Intelligence, {IJCAI-22}},
  publisher = {International Joint Conferences on Artificial Intelligence Organization},
  editor    = {Lud De Raedt},
  pages     = {5841--5842},
  year      = {2022},
  month     = {7},
  note      = {Doctoral Consortium},
  doi       = {10.24963/ijcai.2022/823},
  url       = {https://doi.org/10.24963/ijcai.2022/823},
}
Dynamic Bandits with Temporal Structure · IJCAI 2022