Toward Optimal Solution for the Context-Attentive Bandit Problem
Djallel Bouneffouf, Raphael Feraud, Sohini Upadhyay, Irina Rish, Yasaman Khazaeni
Abstract
In various recommender system applications, from medical diagnosis to dialog systems, due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. In this paper, we analyze and extend an online learning framework known as Context-Attentive Bandit, We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets.
BibTeX
@inproceedings{ijcai2021p481,
title = {Toward Optimal Solution for the Context-Attentive Bandit Problem},
author = {Bouneffouf, Djallel and Feraud, Raphael and Upadhyay, Sohini and Rish, Irina and Khazaeni, Yasaman},
booktitle = {Proceedings of the Thirtieth International Joint Conference on
Artificial Intelligence, {IJCAI-21}},
publisher = {International Joint Conferences on Artificial Intelligence Organization},
editor = {Zhi-Hua Zhou},
pages = {3493--3500},
year = {2021},
month = {8},
note = {Main Track},
doi = {10.24963/ijcai.2021/481},
url = {https://doi.org/10.24963/ijcai.2021/481},
}