2021
Optimal Thompson Sampling strategies for support-aware CVaR bandits
ICML 2021spotlight
In this paper we study a multi-arm bandit problem in which the quality of each arm is measured by the Conditional Value at Risk (CVaR) at some level alpha of the reward distribution. While existing works in this setting mainly focus on Upper Confidence Bound algorithms, we introduce a new Thompson S…