2020
Stochastic Bandits with Delay-Dependent Payoffs
AISTATS 2020poster
Motivated by recommendation problems in music streaming platforms, we propose a nonstationary stochastic bandit model in which the expected reward of an arm depends on the number of rounds that have passed since the arm was last pulled. After proving that finding an optimal policy is NP-hard even wh…