← Search

Cem Kalkanli

1 accepted papers

2021

Batched Thompson Sampling

NeurIPS 2021poster

We introduce a novel anytime batched Thompson sampling policy for multi-armed bandits where the agent observes the rewards of her actions and adjusts her policy only at the end of a small number of batches. We show that this policy simultaneously achieves a problem dependent regret of order $O(\log(…

Cited by 20SourcePDFScholar