Energy Regularized RNNS for solving non-stationary Bandit problems
We consider a Multi-Armed Bandit problem in which the re- wards are non-stationary and are dependent on past actions and potentially on past contexts. At the heart of our method, we employ a recurrent neural network, which models these sequences. In order to balance between exploration and exploitat…