2026
Best-of-three-worlds Analysis for Dueling Bandits with Borda Winner
ICLR 2026poster
The dueling bandits (DB) problem addresses online learning from relative preferences, where the learner queries pairs of arms and receives binary win-loss feedback. Most existing work focuses on designing algorithms for specific stochastic or adversarial environments. Recently, a unified algorithm h…