2023
Langevin Thompson Sampling with Logarithmic Communication: Bandits and Reinforcement Learning
ICML 2023poster
Thompson sampling (TS) is widely used in sequential decision making due to its ease of use and appealing empirical performance. However, many existing analytical and empirical results for TS rely on restrictive assumptions on reward distributions, such as belonging to conjugate families, which limit…