← Search

Samuel McLaughlin Robertson

2 accepted papers

2025

Eluder dimension: localise it!

NeurIPS 2025spotlight

We establish a lower bound on the eluder dimension in generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret bounds. To address this, we introduce a localisation method for the eluder dimension; our analysis immediately recovers and…

Cited by 0SourceScholar
2025

REINFORCE Converges to Optimal Policies with Any Learning Rate

NeurIPS 2025poster

We prove that the classic REINFORCE stochastic policy gradient (SPG) method converges to globally optimal policies in finite-horizon Markov Decision Processes (MDPs) with $\textit{any}$ constant learning rate. To avoid the need for small or decaying learning rates, we introduce two key innovations i…

Cited by 0SourceScholar