2019
Individual Regret in Cooperative Nonstochastic Multi-Armed Bandits
NeurIPS 2019poster
We study agents communicating over an underlying network by exchanging messages, in order to optimize their individual regret in a common nonstochastic multi-armed bandit problem. We derive regret minimization algorithms that guarantee for each agent $v$ an individual expected regret of $\widetilde{…