ICML 2022spotlight28 citations
Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences
Aadirupa Saha, Pierre Gaillard
Abstract
We study the problem of $K$-armed dueling bandit for both stochastic and adversarial environments, where the goal of the learner is to aggregate information through relative preferences of pair of decision points queried in an online sequential manner. We first propose a novel reduction from any (general) dueling bandits to multi-armed bandits which allows us to improve many existing results in dueling bandits. In particular,
BibTeX
@InProceedings{pmlr-v162-saha22a,
title = {Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative Preferences},
author = {Saha, Aadirupa and Gaillard, Pierre},
booktitle = {Proceedings of the 39th International Conference on Machine Learning},
pages = {19011--19026},
year = {2022},
editor = {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
volume = {162},
series = {Proceedings of Machine Learning Research},
month = {17--23 Jul},
publisher = {PMLR},
pdf = {https://proceedings.mlr.press/v162/saha22a/saha22a.pdf},
url = {https://proceedings.mlr.press/v162/saha22a.html},
abstract = {We study the problem of $K$-armed dueling bandit for both stochastic and adversarial environments, where the goal of the learner is to aggregate information through relative preferences of pair of decision points queried in an online sequential manner. We first propose a novel reduction from any (general) dueling bandits to multi-armed bandits which allows us to improve many existing results in dueling bandits. In particular,