← Search

Thang D. Chu

1 accepted papers

2025

REINFORCE Converges to Optimal Policies with Any Learning Rate

NeurIPS 2025poster

We prove that the classic REINFORCE stochastic policy gradient (SPG) method converges to globally optimal policies in finite-horizon Markov Decision Processes (MDPs) with $\textit{any}$ constant learning rate. To avoid the need for small or decaying learning rates, we introduce two key innovations i…

Cited by 0SourceScholar