← Search

Costis Maglaras

1 accepted papers

2019

Thompson Sampling with Information Relaxation Penalties

NeurIPS 2019poster

We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoi…