2021
Online Hyper-Parameter Tuning for the Contextual Bandit
ICASSP 2021accepted
We study here the problem of learning the exploration exploitation trad-off in the contextual bandit problem with linear reward function setting. In the traditional algorithms that solve the contextual bandit problem, the exploration is a parameter that is tuned by the user. However, our proposed al…