2018
PG-TS: Improved Thompson Sampling for Logistic Contextual Bandits
NeurIPS 2018poster
We address the problem of regret minimization in logistic contextual bandits, where a learner decides among sequential actions or arms given their respective contexts to maximize binary rewards. Using a fast inference procedure with Polya-Gamma distributed augmentation variables, we propose an impro…