2025
Data-Driven Upper Confidence Bounds with Near-Optimal Regret for Heavy-Tailed Bandits
AISTATS 2025poster
Stochastic multi-armed bandits (MABs) provide a fundamental reinforcement learning model to study sequential decision making in uncertain environments. The upper confidence bounds (UCB) algorithm gave birth to the renaissance of bandit algorithms, as it achieves near-optimal regret rates under vario…