A near-optimal high-probability swap-Regret upper bound for multi-agent bandits in unknown general-sum games
In this paper, we study a multi-agent bandit problem in an unknown general-sum game repeated for a number of rounds (i.e., learning in a black-box game with bandit feedback), where a set of agents have no information about the underlying game structure and cannot observe each other’s actions and rew…