2023
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
NeurIPS 2023poster
In this paper, we prove state-of-the-art Bayesian regret bounds for Thompson Sampling in reinforcement learning in a multitude of settings. We present a refined analysis of the information ratio, and show an upper bound of order $\widetilde{O}(H\sqrt{d_{l_1}T})$ in the time inhomogeneous reinforceme…