← Search

Hamish Flynn

4 accepted papers

2026

Posterior Sampling Reinforcement Learning with Gaussian Processes for Continuous Control: Sublinear Regret Bounds for Unbounded State Spaces

ICML 2026poster

We analyze the Bayesian regret of the Gaussian process posterior sampling reinforcement learning (GP-PSRL) algorithm. Posterior sampling is an effective heuristic for decision-making under uncertainty that has been used to develop successful algorithms for a variety of continuous control problems. H…

Cited by 0SourceScholar
2023

Improved Algorithms for Stochastic Linear Bandits Using Tail Bounds for Martingale Mixtures

NeurIPS 2023oral

We present improved algorithms with worst-case regret guarantees for the stochastic linear bandit problem. The widely used "optimism in the face of uncertainty" principle reduces a stochastic bandit problem to the construction of a confidence sequence for the unknown reward function. The performance…

Cited by 9SourcePDFScholar