← Search

Tobias Sommer Thune

3 accepted papers

2019

Nonstochastic Multiarmed Bandits with Unrestricted Delays

NeurIPS 2019poster

We investigate multiarmed bandits with delayed feedback, where the delays need neither be identical nor bounded. We first prove that "delayed" Exp3 achieves the $O(\sqrt{(KT + D)\ln K})$ regret bound conjectured by Cesa-Bianchi et al. [2016] in the case of variable, but bounded delays. Here, $K$ is…

Cited by 65SourcePDFScholar