Double-Linear Thompson Sampling for Context-Attentive Bandits
Djallel Bouneffouf, Raphaël Féraud, Sohini Upadhyay, Yasaman Khazaeni, Irina Rish
Abstract
In this paper, we analyze and extend an online learning frame-work known as Context-Attentive Bandit, motivated by various practical applications, from medical diagnosis to dialog systems, where due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets.
BibTeX
@inproceedings{icassp2021_doublelinearthom,
title = {Double-Linear Thompson Sampling for Context-Attentive Bandits},
author = {Djallel Bouneffouf and Raphaël Féraud and Sohini Upadhyay and Yasaman Khazaeni and Irina Rish},
booktitle = {ICASSP 2021},
year = {2021}
}