← Search

Randy Jia

1 accepted papers

2017

Optimistic posterior sampling for reinforcement learning: worst-case regret bounds

NeurIPS 2017poster

We present an algorithm based on posterior sampling (aka Thompson sampling) that achieves near-optimal worst-case regret bounds when the underlying Markov Decision Process (MDP) is communicating with a finite, though unknown, diameter. Our main result is a high probability regret upper bound of $\ti…

Cited by 267SourcePDFScholar