← Search

Priyank Agrawal

4 accepted papers

2025

On the Convergence of Single-Timescale Actor-Critic

NeurIPS 2025poster

We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces. To this end, we introduce an elegant analytical framework for handling complex, coupled recursions inherent in the algo…

Cited by 0SourceScholar
2021

Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration

AAAI 2021technical

This paper studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our $tilde{…

Cited by 24SourcePDFScholar