2016
Dueling Bandits: Beyond Condorcet Winners to General Tournament Solutions
NeurIPS 2016poster
Recent work on deriving $O(\log T)$ anytime regret bounds for stochastic dueling bandit problems has considered mostly Condorcet winners, which do not always exist, and more recently, winners defined by the Copeland set, which do always exist. In this work, we consider a broad notion of winners defi…